Mmastodon TechnologySoftware first seen 1 h ago, last 1 h ago, peak #1
Multi-Token Prediction Boosts RTX 3090 LLM Speed
Original: Originally published on my blog. Enabling MTP on this RTX 3090 raised generation throughput from... # ai # llm # program
A developer reports enabling multi-token prediction (MTP) on an RTX 3090 graphics card raised local LLM generation throughput, while questioning whether the speedup affects coding quality. The write-up, originally published on a personal blog, has drawn attention from AI and open-source software communities interested in getting more performance from consumer GPUs for running large language models locally.
Why now: Hobbyists running local LLMs are keen to learn performance tricks for consumer hardware like the RTX 3090.
RTX 3090NVIDIAmulti-token prediction
Rank over time, top of the chart is #1. 2 snapshots from 1 h ago to 1 h ago.
Evidence
- Originally published on my blog. Enabling MTP on this RTX 3090 raised generation throughput from... # ai # llm # programming # opensource # software # coding # development # engineering # inclusive # community MTP on an RTX 3090: Faster Tokens, but What About Coding Quality? · hackaday@www.urbanmind.net · 5
API: https://socialmediatrends-api.osmike.com/v1/trends/785734