MikeTrendsTrends right now

Mmastodon TechnologySoftware first seen 1 h ago, last 1 h ago, peak #1

Multi-Token Prediction Boosts RTX 3090 LLM Speed

Original: Originally published on my blog. Enabling MTP on this RTX 3090 raised generation throughput from... # ai # llm # program

A developer reports enabling multi-token prediction (MTP) on an RTX 3090 graphics card raised local LLM generation throughput, while questioning whether the speedup affects coding quality. The write-up, originally published on a personal blog, has drawn attention from AI and open-source software communities interested in getting more performance from consumer GPUs for running large language models locally.

Why now: Hobbyists running local LLMs are keen to learn performance tricks for consumer hardware like the RTX 3090.

RTX 3090NVIDIAmulti-token prediction

Open on mastodon →

Rank over time, top of the chart is #1. 2 snapshots from 1 h ago to 1 h ago.

Evidence

API: https://socialmediatrends-api.osmike.com/v1/trends/785734