1M context with 17 GB model in 24 GB VRAM: "for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text"
2.55T2 sourcer/LocalLLaMA
Source record
Published by r/LocalLLaMA (T2 source). The original is at https://www.reddit.com/r/LocalLLaMA/comments/1vkicyd/1m_context_with_17_gb_model_in_24_gb_vram_for_the/.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryA user reports successfully loading a 1M-token context for a ~17GB LLM within 24GB of VRAM and retrieving information from multiple positions in the text, demonstrating large-context inference on consumer-grade GPU memory.
Why it mattersConcrete working numbers for 1M-context inference on a 24GB GPU. Useful reference point for anyone sizing local LLM hardware against long-context workloads.
Cited by
No citations on record.
