The pool, cell by cell
pool snapshot 2026-09-18 · every row is a real community cell with its basis and run count · ranking provisional · a signed run outranks any claim
reading the pool...
pool snapshot 2026-09-18 · every row is a real community cell with its basis and run count · ranking provisional · a signed run outranks any claim
reading the pool...
60 cells shown · 13 measured · 47 reported · MIN_RUNS_MEASURED = 3 · capture or correct a number ->
| # | model | rig | basis | value | n | ctx | bits |
|---|---|---|---|---|---|---|---|
| 1 | stories-llama2-50kchat | AMD Ryzen 7 9800X3D | reported | 36,715.7 tok/s | 1 | 2,048 | 3-bit |
| 2 | models-movedchat | AMD Ryzen 7 9800X3D | reported | 16,740.7 tok/s | 1 | 2,048 | 3-bit |
| 3 | Qwen3.8-27B-NVFP4chat | RTX 5090 32GB | reported | 4,662.6 tok/s | 2 | 4,096 | 4-bit |
| 4 | models-movedchat | RTX 5090 32GB | reported | 4,360.9 tok/s | 1 | 2,048 | 3-bit |
| 5 | diffusiongemma-26B-A4B-it-FP8-dynamicchat | H100 80GB | reported | 1,369.3 tok/s | 1 | 8,192 | 8-bit |
| 6 | Qwen3-0.6B-GGUFchat | RTX PRO 6000 Blackwell 96GB | reported | 956.4 tok/s | 1 | 640 | 4-bit |
| 7 | Llama-3.2-1B-Instruct-GGUFchat | RTX PRO 6000 Blackwell 96GB | reported | 699.9 tok/s | 1 | 4,096 | 8-bit |
| 8 | LFM2.5-1.2B-Instruct-GGUFchat | RTX 3090 24GB | measured | 682.2 tok/s | 4 | 24,576 | 4-bit |
| 9 | gemma-3-1b-itchat | RTX PRO 6000 Blackwell 96GB | reported | 664.3 tok/s | 1 | 640 | 4-bit |
| 10 | NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4chat | RTX 5090 32GB | measured | 582.5 tok/s | 3 | 2,048 | 4-bit |
| 11 | LFM2.5-1.2B-Instruct-GGUFchat | RTX 3090 24GB | measured | 564.4 tok/s | 3 | 24,576 | 6-bit |
| 12 | NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4chat | RTX 5090 32GB | reported | 564.1 tok/s | 2 | 65,536 | 4-bit |
| 13 | diffusiongemma-26B-A4B-itchat | RTX PRO 6000 Blackwell 96GB | reported | 528.7 tok/s | 1 | 8,192 | 16-bit |
| 14 | Qwen3.6-35B-A3B-NVFP4chat | RTX PRO 6000 Blackwell 96GB | reported | 506.2 tok/s | 1 | 8,192 | 4-bit |
| 15 | LFM2.5-1.2B-Instructchat | RX 6750 XT 12GB | reported | 500 tok/s | 1 | 64,000 | 4-bit |
| 16 | NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4chat | RTX PRO 6000 Blackwell 96GB | measured | 494.7 tok/s | 4 | 32,768 | 4-bit |
| 17 | Llama-3.2-3B-Instructchat | RTX PRO 6000 Blackwell 96GB | reported | 489.2 tok/s | 1 | 640 | 4-bit |
| 18 | Llama-3.2-1B-Instructchat | RTX 5080 16GB | reported | 448.1 tok/s | 1 | 4,096 | 8-bit |
| 19 | GLM-5.3-Flashchat | RTX PRO 6000 Blackwell 96GB ×2 | measured | 443.1 tok/s | 9 | 8,192 | 4-bit |
| 20 | Qwen3.8-27B-QUASAR-NVFP4chat | Tesla V100 32GB ×4 | reported | 427.3 tok/s | 1 | 262,144 | 4-bit |
| 21 | LFM2.5-8B-A1Bchat | RTX 5070 Ti 16GB | reported | 402.3 tok/s | 1 | 4,096 | 4-bit |
| 22 | gemma-4-26B-A4B-it-GGUFchat | RTX 5090 32GB | reported | 395 tok/s | 2 | 32,768 | 4-bit |
| 23 | LFM2-24B-A2B-GGUFchat | RTX 5090 32GB | reported | 393.2 tok/s | 1 | 2,048 | 4-bit |
| 24 | Qwen3.8-27Bchat | Tesla V100 16GB ×4 | reported | 391 tok/s | 1 | 32,768 | 4-bit |
| 25 | NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16chat | H100 80GB | reported | 381.2 tok/s | 2 | 32,768 | 16-bit |
| 26 | Qwen3.5-9Bchat | Radeon AI Pro R9700 32GB | reported | 371.8 tok/s | 1 | 4,096 | 4-bit |
| 27 | Ling-3.0-tinychat | RTX PRO 6000 Blackwell 96GB | reported | 369.1 tok/s | 1 | 640 | 4-bit |
| 28 | gemma-4-26B-A4B-it-AWQ-4bitchat | RTX 5090 32GB | measured | 369 tok/s | 4 | 2,048 | 4-bit |
| 29 | Unlimited-OCRchat | RTX 3090 24GB | reported | 365.5 tok/s | 1 | 2,048 | 16-bit |
| 30 | NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUFchat | RTX 3090 24GB | reported | 353.7 tok/s | 1 | 8,192 | 4-bit |
| 31 | Qwen3.6-35B-A3B-NVFP4chat | RTX 5090 32GB ×2 | reported | 340.8 tok/s | 1 | 32,768 | 4-bit |
| 32 | NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16chat | H200 SXM 141GB | reported | 337.6 tok/s | 1 | 32,768 | 16-bit |
| 33 | Qwen3.6-35B-A3B-FP8chat | H100 80GB | reported | 337.1 tok/s | 2 | 32,768 | 8-bit |
| 34 | Qwen3.5-9Bchat | RX 7900 XTX 24GB | measured | 337 tok/s | 7 | 4,096 | 4-bit |
| 35 | Qwen3.6-35B-A3B-NVFP4chat | RTX PRO 6000 Blackwell 96GB | measured | 330.3 tok/s | 3 | 262,144 | 4-bit |
| 36 | LFM2.5-8B-A1B-GGUFchat | RTX 3080 12GB | reported | 328.4 tok/s | 1 | 128,000 | 4-bit |
| 37 | Qwen3.6-35B-A3B-NVFP4chat | RTX PRO 6000 Blackwell 96GB ×4 | reported | 325.9 tok/s | 1 | 2,048 | 4-bit |
| 38 | gemma-3-1b-it-GGUFchat | RTX A5000 24GB | reported | 325.2 tok/s | 2 | 4,096 | 4-bit |
| 39 | Qwen3.6-35B-A3B-GGUFchat | RTX 5090 32GB | reported | 324.4 tok/s | 2 | 32,768 | 4-bit |
| 40 | Qwen2.5-1.5B-Instruct-GGUFchat | RTX A5000 24GB | reported | 320.6 tok/s | 2 | 4,096 | 4-bit |
| 41 | DeepSeek-V4.1-Flash-UNCENSORED-FP8chat | RTX PRO 6000 Blackwell Max-Q Workstation Edition 96GB ×4 | reported | 316.6 tok/s | 1 | 524,288 | 8-bit |
| 42 | Qwen3-30B-A3B-Instruct-2507-GGUFchat | RTX 5090 32GB | reported | 312.9 tok/s | 1 | 2,048 | 4-bit |
| 43 | Qwen3-0.6Bchat | RTX 3060 12GB | reported | 310.4 tok/s | 1 | 4,608 | 8-bit |
| 44 | DeepSeek-V4-Flash-Abliterated-DSparkchat | RTX PRO 6000 Blackwell 96GB ×4 | reported | 309 tok/s | 2 | 2,048 | 8-bit |
| 45 | DeepSeek-V4-Flash-DSparkchat | RTX PRO 6000 Blackwell 96GB ×4 | measured | 308.2 tok/s | 3 | 2,048 | 8-bit |
| 46 | LFM2.5-8B-A1Bchat | RTX 5070 Ti 16GB | reported | 307.4 tok/s | 2 | 4,096 | 8-bit |
| 47 | gpt-oss-20bchat | RTX 5090 32GB | reported | 307.2 tok/s | 1 | 2,048 | 4-bit |
| 48 | LFM2.5-2.6B-GGUFchat | RTX 3090 24GB | reported | 305.6 tok/s | 1 | 24,576 | 4-bit |
| 49 | DeepSeek-V4.1-Flashchat | RTX PRO 6000 Blackwell Max-Q Workstation Edition 96GB ×4 | measured | 303.7 tok/s | 6 | 524,288 | 8-bit |
| 50 | NVIDIA-Nemotron-3-Nano-30B-A3B-BF16chat | RTX 5090 32GB | reported | 299.4 tok/s | 2 | 8,192 | 4-bit |
| 51 | DeepSeek-V4-Flash-0731chat | RTX PRO 6000 Blackwell 96GB ×4 | measured | 296.7 tok/s | 5 | 2,048 | 8-bit |
| 52 | Laguna-XS-2.1chat | RTX 3090 24GB | reported | 296 tok/s | 1 | 2,048 | 4-bit |
| 53 | Qwen3.8-27B-NVFP4-BF16-LMHeadchat | RTX PRO 6000 Blackwell 96GB | reported | 296 tok/s | 1 | 262,144 | 4-bit |
| 54 | Qwen3.8-27Bchat | H100 NVL 94GB ×2 | reported | 293.3 tok/s | 1 | 2,048 | 8-bit |
| 55 | Qwen3-Coder-30B-A3B-Instruct-AWQcode | RTX 5090 32GB | reported | 287.7 tok/s | 2 | 2,048 | 4-bit |
| 56 | NVIDIA-Nemotron-3-Nano-30B-A3B-GGUFchat | RTX 5090 32GB | reported | 286.8 tok/s | 1 | 2,048 | 4-bit |
| 57 | Qwen3.5-27Bchat | Radeon AI Pro R9700 34GB | reported | 286.6 tok/s | 1 | 4,096 | 4-bit |
| 58 | Qwen3.8-27Bchat | Radeon AI Pro R9700 32GB | measured | 285.5 tok/s | 7 | 98,304 | 4-bit |
| 59 | gemma-4-26B-A4B-itchat | RTX 5090 32GB | measured | 282 tok/s | 7 | 2,048 | 8-bit |
| 60 | gemma-3-1b-it-GGUFchat | RTX A5000 24GB | reported | 281.9 tok/s | 2 | 4,096 | 8-bit |
Rows are community cells (medians), ranked provisionally. A signed run outranks any claim.