Local LLM
Language models on hardware you already own: quantization trade-offs, context length limits, and tokens per second you can reproduce.
Tokens per second only means something next to the quantization, the context length and the runtime version it was measured with. Every record in this cluster carries all four, so a number measured here can be compared with a number measured on your own machine.
This cluster opens in week 7 of the plan. Until then, the video cluster and the compatibility database carry the measurements.