| 1 | Tinystories lay8 hs512 hd8 33mReported as ivnle/tinystories-lay8-hs512-hd8-33M · llama · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 2,068.0 TG128 | 190,107 PP512 | arki05 | 17 Sep 2026 | |
| 2 | Tinystories lay8 hs512 hd8 33mReported as RichardErkhov/ivnle_-_tinystories-lay8-hs512-hd8-33M-gguf · llama Q4_0 | llama.cpp | Apple M5 Pro BLAS + Metal | 1,443.3 TG128 | 124,181 PP512 | arki05 | 17 Sep 2026 | |
| 3 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M5 Max Metal | 708.3 TG128 | 34,136 PP512 | prabod | 10 Sep 2026 | |
| 4 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8 Q8 | BaseRT | Apple M5 Max Metal | 549.5 TG128 | 33,088 PP512 | prabod | 10 Sep 2026 | |
| 5 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 516.8 TG128 | 20,778 PP512 | arki05 | 11 Sep 2026 | |
| 6 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 508.8 TG128 | 26,137 PP512 | arki05 | 11 Sep 2026 | |
| 7 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 479.4 TG128 | 26,346 PP512 | arki05 | 11 Sep 2026 | |
| 8 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | Apple M5 Max BLAS + Metal | 439.1 TG128 | 24,365 PP512 | prabod | 10 Sep 2026 | |
| 9 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | Apple M5 Max BLAS + Metal | 379.0 TG128 | 24,983 PP512 | prabod | 10 Sep 2026 | |
| 10 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT ROCm | 368.6 TG128 | 23,972 PP512 | arki05 | 11 Sep 2026 | |
| 11 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q4_K_M | llama.cpp | Apple M5 Pro BLAS + Metal | 358.3 TG128 | 14,552 PP512 | arki05 | 11 Sep 2026 | |
| 12 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q8 Q8 | BaseRT | Apple M5 Pro Metal | 350.2 TG128 | 20,537 PP512 | arki05 | 11 Sep 2026 | |
| 13 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | AMD Radeon RX 7900 XT ROCm | 313.7 TG128 | 24,691 PP512 | arki05 | 11 Sep 2026 | |
| 14 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 293.1 TG128 | 15,066 PP512 | arki05 | 11 Sep 2026 | |
| 15 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen3 Q8_0 | llama.cpp | Apple M5 Pro BLAS + Metal | 283.0 TG128 | 14,942 PP512 | arki05 | 11 Sep 2026 | |
| 16 | Qwen3 1.7BReported as Qwen/Qwen3-1.7B · qwen · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 235.4 TG128 | 8,040 PP512 | arki05 | 11 Sep 2026 | |
| 17 | Llama 3.2 3B InstructReported as meta-llama/Llama-3.2-3B-Instruct · llama Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 229.7 TG128 | 5,737 PP512 | arki05 | 11 Sep 2026 | |
| 18 | Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 229.1 TG128 | 11,673 PP512 | arki05 | 11 Sep 2026 | |
| 19 | Qwen3.5 2B BaseReported as Qwen/Qwen3.5-2B-Base · qwen35 · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 218.2 TG128 | 2,645 PP512 | arki05 | 11 Sep 2026 | |
| 20 | Qwen3.5 2BReported as Qwen/Qwen3.5-2B · qwen35 · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 217.7 TG128 | 2,641 PP512 | arki05 | 11 Sep 2026 | |
| 21 | Gemma 3 1B ITReported as google/gemma-3-1b-it · gemma3 · default-q8 Q8 | BaseRT | Apple M5 Pro Metal | 210.3 TG128 | 15,061 PP512 | arki05 | 11 Sep 2026 | |
| 22 | Llama 3.2 1B InstructReported as meta-llama/Llama-3.2-1B-Instruct · llama · default-q8 Q4 | BaseRT | Apple M5 Pro Metal | 205.3 TG128 | 12,067 PP512 | arki05 | 11 Sep 2026 | |
| 23 | Qwen3 30B A3B Instruct 2507Reported as Qwen/Qwen3-30B-A3B-Instruct-2507 · qwen3moe Q2_K | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 198.1 TG128 | 2,837 PP512 | arki05 | 11 Sep 2026 | |
| 24 | Gemma 4 E2B ITReported as google/gemma-4-E2B-it · gemma4 · default-q4 Q4 | BaseRT | Apple M5 Max Metal | 197.5 TG128 | 20,099 PP512 | lukas | 15 Sep 2026 | |
| 25 | GPT-OSS 20BReported as openai/gpt-oss-20b · gpt-oss Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 195.6 TG128 | 3,292 PP512 | arki05 | 11 Sep 2026 | |
| 26 | Qwen3 30B A3B Instruct 2507Reported as Qwen/Qwen3-30B-A3B-Instruct-2507 · qwen3moe Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 186.7 TG128 | 2,868 PP512 | arki05 | 11 Sep 2026 | |
| 27 | Qwen3 4B Instruct 2507Reported as Qwen/Qwen3-4B-Instruct-2507 · qwen3 Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 185.1 TG128 | 4,793 PP512 | arki05 | 11 Sep 2026 | |
| 28 | Qwen3 30B A3B Instruct 2507Reported as Qwen/Qwen3-30B-A3B-Instruct-2507 · qwen3moe Q3_K_M | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 184.5 TG128 | 2,474 PP512 | arki05 | 11 Sep 2026 | |
| 29 | Nemotron 3 Nano 30B A3BReported as nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 · nemotron_h_moe · default-q4 Q4 | BaseRT | Apple M5 Max Metal | 184.4 TG128 | 4,984 PP512 | lukas | 11 Sep 2026 | |
| 30 | Qwen2.5 Coder 1.5B InstructReported as Qwen/Qwen2.5-Coder-1.5B-Instruct · qwen2 Q4_K_M | llama.cpp | NVIDIA GeForce RTX 4060 Laptop GPU CUDA | 178.8 TG128 | 9,465 PP512 | sarthak247 | 11 Sep 2026 | |
| 31 | Qwen3 0.6BReported as Qwen/Qwen3-0.6B · qwen · default-q4 Q4 | BaseRT | Apple M5 Metal | 172.7 TG128 | 4,997 PP512 | skogul97 | 11 Sep 2026 | |
| 32 | Llama 3.2 3B InstructReported as meta-llama/Llama-3.2-3B-Instruct · llama Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT ROCm | 172.0 TG128 | 6,980 PP512 | arki05 | 11 Sep 2026 | |
| 33 | GPT-OSS 20BReported as openai/gpt-oss-20b · gpt_oss · default-q4 Q4 | BaseRT | Apple M5 Max Metal | 169.0 TG128 | 2,179 PP512 | lukas | 11 Sep 2026 | |
| 34 | Llama 3.2 3B InstructReported as meta-llama/Llama-3.2-3B-Instruct · llama Q8_0 | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 167.5 TG128 | 5,846 PP512 | arki05 | 11 Sep 2026 | |
| 35 | GPT-OSS 20BReported as openai/gpt-oss-20b · gpt-oss Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT ROCm | 160.1 TG128 | 3,502 PP512 | arki05 | 11 Sep 2026 | |
| 36 | Qwen 0.6B Coder (XformAI)Reported as XformAI-india/qwen-0.6b-coder · qwen3 Q2_K | llama.cpp | Apple M1 Pro BLAS + Metal | 151.6 TG128 | 2,620 PP512 | lukas | 10 Sep 2026 | |
| 37 | Qwen3 8BReported as Qwen/Qwen3-8B · qwen3 Q2_K | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 150.7 TG128 | 2,563 PP512 | arki05 | 11 Sep 2026 | |
| 38 | Qwen3 4B Instruct 2507Reported as Qwen/Qwen3-4B-Instruct-2507 · qwen3 Q4_K_M | llama.cpp | AMD Radeon RX 7900 XT ROCm | 147.8 TG128 | 5,364 PP512 | arki05 | 11 Sep 2026 | |
| 39 | GPT-OSS 20BReported as openai/gpt-oss-20b · gpt-oss F16 | llama.cpp | AMD Radeon RX 7900 XT (RADV NAVI31) Vulkan | 146.8 TG128 | 3,262 PP512 | arki05 | 11 Sep 2026 | |
| 40 | Gemma 4 E2B ITReported as google/gemma-4-E2B-it · gemma4 · default-q4 Q4 | BaseRT | Apple M5 Pro Metal | 144.6 TG128 | 13,240 PP512 | arki05 | 11 Sep 2026 | |