Every open-source LLM, ranked daily.
Every open-weight model worth knowing, ranked two ways: by human votes (LMArena Elo) and by benchmarks (Artificial Analysis), so the newest models show up the day they launch instead of weeks later. With the exact RAM each quant needs (real file sizes) and where to download it. As of July 27, 2026, the human-vote leader is glm-5.1.
Updated July 27, 2026 · Sources: LMArena style-controlled Elo + Artificial Analysis · quant sizes from real GGUF files · 165 rated + 10 new
-
1 glm-5.1 Zhipu · MIT 355B 206GB +
- Chat Elo (human votes)
- 1470(1465-1474)
- AA Intelligence (benchmark)
- 40.2 · code 55.8
- WebDev Elo (coding)
- 1520
- Arena votes
- 30,726
- Size
- 355B (32B active)
- Context
- 205K
- Multimodal
- text only
- Smallest rig that fits
- 256GB
Zhipu's flagship. Excellent all-rounder, strong coding, clean MIT license. Best for: Coding and general agent work.
Benchmarks by category (LMArena Elo)
- Coding
- 1520
- Math
- 1480
- Creative writing
- 1454
- Instruction following
- 1464
- Hard prompts
- 1492
- Long queries
- 1484
- Multi-turn chat
- 1483
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 206GB UD-IQ2_XXS 221GB UD-IQ2_M 236GB UD-Q2_K_XL 252GB UD-Q3_K_XL 340GB UD-Q4_K_XL 466GB UD-Q5_K_XL 560GB Q8_0 801GB -
2 glm-5.2 (max) Zhipu · MIT ? ? +
- Chat Elo (human votes)
- 1469(1463-1475)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1588
- Arena votes
- 18,017
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1509
- Math
- 1474
- Creative writing
- 1446
- Instruction following
- 1463
- Hard prompts
- 1490
- Long queries
- 1479
- Multi-turn chat
- 1471
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
3 mimo-v2.5-pro Xiaomi · MIT ? 304GB +
- Chat Elo (human votes)
- 1467(1462-1471)
- AA Intelligence (benchmark)
- 42.2 · code 60.2
- WebDev Elo (coding)
- 1476
- Arena votes
- 42,195
- Size
- unconfirmed
- Context
- 1M
- Multimodal
- text only
- Smallest rig that fits
- 512GB Mac Studio
Xiaomi's flagship MiMo. Strong scores; confirm size on the model card. Best for: General use (verify specs).
Benchmarks by category (LMArena Elo)
- Coding
- 1520
- Math
- 1475
- Creative writing
- 1433
- Instruction following
- 1470
- Hard prompts
- 1495
- Long queries
- 1488
- Multi-turn chat
- 1477
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 304GB UD-IQ2_XXS 317GB UD-IQ2_M 317GB UD-Q2_K_XL 338GB UD-Q3_K_XL 460GB UD-Q4_K_XL 631GB UD-Q5_K_XL 759GB Q8_0 1088GB -
4 kimi-k2.6 Moonshot AI · Modified MIT 1000B 340GB +
- Chat Elo (human votes)
- 1461(1456-1465)
- AA Intelligence (benchmark)
- 44.2 · code 61.8
- WebDev Elo (coding)
- 1510
- Arena votes
- 37,686
- Size
- 1000B (32B active)
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 512GB Mac Studio
Frontier-class open reasoning. Huge MoE, so it needs serious memory, but on a big-RAM Mac it flies. Best for: Top-end local reasoning and agents.
Benchmarks by category (LMArena Elo)
- Coding
- 1515
- Math
- 1480
- Creative writing
- 1430
- Instruction following
- 1455
- Hard prompts
- 1485
- Long queries
- 1476
- Multi-turn chat
- 1460
RAM per quant (real file sizes, via unsloth)
UD-Q2_K_XL 340GB UD-Q4_K_XL 584GB -
5 glm-5 Zhipu · MIT 355B 204GB +
- Chat Elo (human votes)
- 1457(1452-1461)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1435
- Arena votes
- 27,785
- Size
- 355B (32B active)
- Context
- 205K
- Multimodal
- text only
- Smallest rig that fits
- 256GB
The prior GLM-5. Still elite, often cheaper to run than 5.1. Best for: General use, coding.
Benchmarks by category (LMArena Elo)
- Coding
- 1497
- Math
- 1443
- Creative writing
- 1445
- Instruction following
- 1447
- Hard prompts
- 1478
- Long queries
- 1470
- Multi-turn chat
- 1472
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 204GB UD-IQ1_M 224GB UD-IQ2_XXS 241GB UD-IQ2_M 255GB Q2_K 276GB UD-Q2_K_XL 281GB Q3_K_M 360GB UD-Q3_K_XL 332GB IQ4_XS 403GB Q4_K_M 456GB UD-Q4_K_XL 431GB Q5_K_M 535GB UD-Q5_K_XL 536GB Q6_K 619GB Q8_0 801GB -
6 deepseek-v4-pro DeepSeek · MIT 671B 272GB est +
- Chat Elo (human votes)
- 1457(1452-1461)
- AA Intelligence (benchmark)
- 44.3 · code 59.4
- WebDev Elo (coding)
- 1447
- Arena votes
- 45,278
- Size
- 671B (37B active)
- Context
- 1M
- Multimodal
- text only
- Smallest rig that fits
- 512GB Mac Studio
DeepSeek's big MoE. Exceptional at code and math, permissive MIT. Best for: Coding, math, large-context work.
Benchmarks by category (LMArena Elo)
- Coding
- 1501
- Math
- 1445
- Creative writing
- 1442
- Instruction following
- 1453
- Hard prompts
- 1480
- Long queries
- 1473
- Multi-turn chat
- 1472
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~373 GB at Q4, ~272 GB at an aggressive dynamic quant.
-
7 deepseek-v4-pro-thinking DeepSeek · MIT 671B 272GB est +
- Chat Elo (human votes)
- 1455(1451-1460)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1464
- Arena votes
- 43,079
- Size
- 671B (37B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 512GB Mac Studio
Reasoning mode of V4 Pro. Best DeepSeek for hard, multi-step problems. Best for: Hard reasoning and coding.
Benchmarks by category (LMArena Elo)
- Coding
- 1490
- Math
- 1467
- Creative writing
- 1443
- Instruction following
- 1447
- Hard prompts
- 1475
- Long queries
- 1466
- Multi-turn chat
- 1456
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~373 GB at Q4, ~272 GB at an aggressive dynamic quant.
-
8 gemma-4-31b Google · Apache 2.0 31B 9GB +
- Chat Elo (human votes)
- 1451(1443-1458)
- AA Intelligence (benchmark)
- 29.4 · code 43.4
- WebDev Elo (coding)
- 1364
- Arena votes
- 5,880
- Size
- 31B
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 16GB
Google's dense 31B. The sweet spot for a single 24GB GPU. Punches above its size. Best for: Single-GPU and laptops.
Benchmarks by category (LMArena Elo)
- Coding
- 1498
- Math
- 1470
- Creative writing
- 1421
- Instruction following
- 1452
- Hard prompts
- 1473
- Long queries
- 1467
- Multi-turn chat
- 1464
RAM per quant (real file sizes, via unsloth)
UD-IQ2_XXS 9GB UD-IQ2_M 11GB UD-Q2_K_XL 12GB Q3_K_M 15GB UD-Q3_K_XL 15GB IQ4_XS 16GB Q4_K_M 18GB UD-Q4_K_XL 19GB Q5_K_M 22GB UD-Q5_K_XL 22GB Q6_K 25GB Q8_0 33GB -
9 kimi-k2.5-thinking Moonshot AI · Modified MIT 1000B 276GB +
- Chat Elo (human votes)
- 1450(1446-1453)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1437
- Arena votes
- 62,688
- Size
- 1000B (32B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 512GB Mac Studio
The deliberate, chain-of-thought Kimi. Slower, stronger on hard problems. Best for: Complex reasoning, planning.
Benchmarks by category (LMArena Elo)
- Coding
- 1502
- Math
- 1472
- Creative writing
- 1424
- Instruction following
- 1440
- Hard prompts
- 1471
- Long queries
- 1458
- Multi-turn chat
- 1452
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 276GB UD-IQ1_M 301GB UD-IQ2_XXS 327GB UD-IQ2_M 345GB Q2_K 374GB UD-Q2_K_XL 375GB Q3_K_M 490GB UD-Q3_K_XL 490GB IQ4_XS 547GB Q4_K_M 621GB UD-Q4_K_XL 622GB Q5_K_M 729GB UD-Q5_K_XL 731GB Q6_K 843GB Q8_0 1091GB -
10 inkling thinky · Apache 2.0 ? 270GB +
- Chat Elo (human votes)
- 1445(1437-1453)
- AA Intelligence (benchmark)
- 40.7 · code 52.1
- WebDev Elo (coding)
- 1418
- Arena votes
- 5,386
- Size
- unconfirmed
- Context
- 1M
- Multimodal
- yes (images)
- Smallest rig that fits
- 512GB Mac Studio
Benchmarks by category (LMArena Elo)
- Coding
- 1498
- Math
- 1476
- Creative writing
- 1395
- Instruction following
- 1428
- Hard prompts
- 1470
- Long queries
- 1448
- Multi-turn chat
- 1454
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 270GB UD-IQ1_M 285GB UD-Q2_K_XL 317GB UD-Q3_K_XL 433GB UD-Q4_K_XL 587GB Q8_0 857GB -
11 minimax-m3 minimax · MiniMax Community License ? 128GB +
- Chat Elo (human votes)
- 1444(1439-1449)
- AA Intelligence (benchmark)
- 44.4 · code 58.6
- WebDev Elo (coding)
- 1493
- Arena votes
- 28,093
- Size
- unconfirmed
- Context
- 1M
- Multimodal
- yes (images)
- Smallest rig that fits
- 128GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1498
- Math
- 1443
- Creative writing
- 1407
- Instruction following
- 1437
- Hard prompts
- 1464
- Long queries
- 1453
- Multi-turn chat
- 1452
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 128GB UD-IQ2_XXS 134GB UD-IQ2_M 134GB UD-Q2_K_XL 143GB UD-Q3_K_XL 195GB UD-Q4_K_XL 265GB UD-Q5_K_XL 318GB Q8_0 453GB -
12 qwen3.5-397b-a17b Alibaba · Apache 2.0 397B 107GB +
- Chat Elo (human votes)
- 1442(1438-1446)
- AA Intelligence (benchmark)
- 33.7 · code 48.2
- WebDev Elo (coding)
- 1400
- Arena votes
- 58,157
- Size
- 397B (17B active)
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 128GB Mac
Alibaba's big sparse MoE: 397B total but only 17B active, so it runs faster than its size suggests. Apache 2.0. Best for: Best license, multilingual, agents.
Benchmarks by category (LMArena Elo)
- Coding
- 1492
- Math
- 1446
- Creative writing
- 1409
- Instruction following
- 1433
- Hard prompts
- 1463
- Long queries
- 1455
- Multi-turn chat
- 1451
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 107GB UD-IQ2_XXS 115GB UD-IQ2_M 123GB Q3_K_M 177GB UD-Q3_K_XL 179GB Q4_K_M 244GB UD-Q4_K_XL 245GB Q5_K_M 294GB UD-Q5_K_XL 295GB Q6_K 327GB Q8_0 422GB -
13 glm-4.7 Zhipu · MIT 355B 97GB +
- Chat Elo (human votes)
- 1442(1436-1448)
- AA Intelligence (benchmark)
- 33.7 · code 45.3
- WebDev Elo (coding)
- 1433
- Arena votes
- 12,092
- Size
- 355B (32B active)
- Context
- 205K
- Multimodal
- text only
- Smallest rig that fits
- 128GB Mac
Last-gen GLM. A proven, stable workhorse if you want maturity over bleeding edge. Best for: Reliable daily driver.
Benchmarks by category (LMArena Elo)
- Coding
- 1485
- Math
- 1428
- Creative writing
- 1405
- Instruction following
- 1428
- Hard prompts
- 1463
- Long queries
- 1452
- Multi-turn chat
- 1459
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 97GB UD-IQ1_M 108GB UD-IQ2_XXS 116GB UD-IQ2_M 122GB Q2_K 131GB UD-Q2_K_XL 135GB Q3_K_M 171GB UD-Q3_K_XL 159GB IQ4_XS 192GB Q4_K_M 216GB UD-Q4_K_XL 205GB Q5_K_M 254GB UD-Q5_K_XL 254GB Q6_K 294GB Q8_0 381GB -
14 deepseek-v4-flash-thinking DeepSeek · MIT 200B 84GB est +
- Chat Elo (human votes)
- 1439(1434-1443)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 44,822
- Size
- 200B (20B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 96GB
The lighter, faster DeepSeek. Easier to run, still a strong reasoner. Best for: Reasoning on mid-range hardware.
Benchmarks by category (LMArena Elo)
- Coding
- 1481
- Math
- 1440
- Creative writing
- 1406
- Instruction following
- 1436
- Hard prompts
- 1460
- Long queries
- 1451
- Multi-turn chat
- 1448
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~114 GB at Q4, ~84 GB at an aggressive dynamic quant.
-
15 gemma-4-26b-a4b Google · Apache 2.0 26B 10GB +
- Chat Elo (human votes)
- 1438(1430-1446)
- AA Intelligence (benchmark)
- 25.7 · code 39.3
- WebDev Elo (coding)
- 1366
- Arena votes
- 5,798
- Size
- 26B (4B active)
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1480
- Math
- 1466
- Creative writing
- 1402
- Instruction following
- 1439
- Hard prompts
- 1461
- Long queries
- 1448
- Multi-turn chat
- 1446
RAM per quant (real file sizes, via unsloth)
UD-IQ2_XXS 10GB UD-IQ2_M 10GB UD-Q2_K_XL 11GB UD-Q3_K_XL 13GB UD-Q4_K_XL 17GB UD-Q5_K_XL 21GB Q8_0 27GB -
16 deepseek-v4-flash DeepSeek · MIT ? 83GB +
- Chat Elo (human votes)
- 1436(1431-1440)
- AA Intelligence (benchmark)
- 40.3 · code 56.2
- WebDev Elo (coding)
- n/a
- Arena votes
- 45,050
- Size
- unconfirmed
- Context
- 1M
- Multimodal
- text only
- Smallest rig that fits
- 96GB
Benchmarks by category (LMArena Elo)
- Coding
- 1482
- Math
- 1426
- Creative writing
- 1408
- Instruction following
- 1428
- Hard prompts
- 1459
- Long queries
- 1449
- Multi-turn chat
- 1453
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 83GB UD-IQ1_M 87GB UD-IQ2_XXS 91GB UD-IQ2_M 91GB UD-Q2_K_XL 97GB UD-Q3_K_XL 129GB UD-Q4_K_XL 155GB -
17 mimo-v2.5 Xiaomi · MIT ? 93GB +
- Chat Elo (human votes)
- 1433(1429-1437)
- AA Intelligence (benchmark)
- 37.2 · code 56.8
- WebDev Elo (coding)
- 1435
- Arena votes
- 43,091
- Size
- unconfirmed
- Context
- 1M
- Multimodal
- yes (images)
- Smallest rig that fits
- 96GB
The standard MiMo 2.5. Lighter than Pro; check the card for exact size. Best for: General use (verify specs).
Benchmarks by category (LMArena Elo)
- Coding
- 1491
- Math
- 1441
- Creative writing
- 1392
- Instruction following
- 1431
- Hard prompts
- 1461
- Long queries
- 1451
- Multi-turn chat
- 1449
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 93GB UD-IQ2_XXS 96GB UD-IQ2_M 97GB UD-Q2_K_XL 103GB UD-Q3_K_XL 140GB UD-Q4_K_XL 192GB UD-Q5_K_XL 231GB Q8_0 329GB -
18 kimi-k2.5-instant Moonshot AI · Modified MIT 1000B 276GB +
- Chat Elo (human votes)
- 1431(1425-1438)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1405
- Arena votes
- 8,178
- Size
- 1000B (32B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 512GB Mac Studio
Kimi tuned for speed over deep thinking. Snappy for chat and tools. Best for: Fast local assistant.
Benchmarks by category (LMArena Elo)
- Coding
- 1504
- Math
- 1441
- Creative writing
- 1390
- Instruction following
- 1435
- Hard prompts
- 1461
- Long queries
- 1445
- Multi-turn chat
- 1439
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 276GB UD-IQ1_M 301GB UD-IQ2_XXS 327GB UD-IQ2_M 345GB Q2_K 374GB UD-Q2_K_XL 375GB Q3_K_M 490GB UD-Q3_K_XL 490GB IQ4_XS 547GB Q4_K_M 621GB UD-Q4_K_XL 622GB Q5_K_M 729GB UD-Q5_K_XL 731GB Q6_K 843GB Q8_0 1091GB -
19 kimi-k2-thinking-turbo Moonshot AI · Modified MIT 1000B 280GB +
- Chat Elo (human votes)
- 1430(1427-1433)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1321
- Arena votes
- 61,950
- Size
- 1000B (32B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 512GB Mac Studio
Faster thinking variant of the previous Kimi gen. Still very capable. Best for: Reasoning on a tighter time budget.
Benchmarks by category (LMArena Elo)
- Coding
- 1486
- Math
- 1438
- Creative writing
- 1392
- Instruction following
- 1417
- Hard prompts
- 1453
- Long queries
- 1433
- Multi-turn chat
- 1433
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 280GB UD-IQ1_M 304GB UD-IQ2_XXS 329GB UD-IQ2_M 347GB Q2_K 373GB UD-Q2_K_XL 382GB Q3_K_M 489GB UD-Q3_K_XL 452GB IQ4_XS 547GB Q4_K_M 621GB UD-Q4_K_XL 587GB Q5_K_M 728GB UD-Q5_K_XL 731GB Q6_K 843GB Q8_0 1091GB -
20 mistral-medium-3.5 mistral · Modified MIT ? ? +
- Chat Elo (human votes)
- 1427(1421-1434)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1267
- Arena votes
- 11,019
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1479
- Math
- 1433
- Creative writing
- 1395
- Instruction following
- 1421
- Hard prompts
- 1446
- Long queries
- 1431
- Multi-turn chat
- 1432
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
21 nvidia-nemotron-3-ultra-550b-a55b-nvfp4 NVIDIA · OpenMDW-1.1 550B 224GB est +
- Chat Elo (human votes)
- 1425(1418-1432)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 10,562
- Size
- 550B (55B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 256GB
Benchmarks by category (LMArena Elo)
- Coding
- 1473
- Math
- 1446
- Creative writing
- 1376
- Instruction following
- 1405
- Hard prompts
- 1442
- Long queries
- 1433
- Multi-turn chat
- 1398
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~307 GB at Q4, ~224 GB at an aggressive dynamic quant.
-
22 glm-4.6 Zhipu · MIT ? 97GB +
- Chat Elo (human votes)
- 1425(1421-1429)
- AA Intelligence (benchmark)
- 28.7 · code 45.8
- WebDev Elo (coding)
- 1339
- Arena votes
- 35,608
- Size
- unconfirmed
- Context
- 205K
- Multimodal
- text only
- Smallest rig that fits
- 128GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1459
- Math
- 1420
- Creative writing
- 1402
- Instruction following
- 1415
- Hard prompts
- 1442
- Long queries
- 1433
- Multi-turn chat
- 1421
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 97GB UD-IQ1_M 107GB UD-IQ2_XXS 115GB UD-IQ2_M 122GB Q2_K 131GB UD-Q2_K_XL 135GB Q3_K_M 171GB UD-Q3_K_XL 158GB IQ4_XS 191GB Q4_K_M 216GB UD-Q4_K_XL 204GB Q5_K_M 253GB UD-Q5_K_XL 252GB Q6_K 293GB Q8_0 379GB -
23 deepseek-v3.2 DeepSeek · MIT ? 184GB +
- Chat Elo (human votes)
- 1425(1421-1428)
- AA Intelligence (benchmark)
- 32.0 · code 44.2
- WebDev Elo (coding)
- 1323
- Arena votes
- 47,206
- Size
- unconfirmed
- Context
- 164K
- Multimodal
- text only
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1470
- Math
- 1429
- Creative writing
- 1401
- Instruction following
- 1421
- Hard prompts
- 1447
- Long queries
- 1442
- Multi-turn chat
- 1428
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 184GB UD-IQ1_M 199GB UD-IQ2_XXS 217GB UD-IQ2_M 228GB Q2_K 245GB UD-Q2_K_XL 247GB Q3_K_M 320GB UD-Q3_K_XL 321GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 408GB Q5_K_M 476GB UD-Q5_K_XL 482GB Q6_K 551GB Q8_0 713GB -
24 deepseek-v3.2-exp-thinking DeepSeek · MIT ? ? +
- Chat Elo (human votes)
- 1425(1418-1431)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 9,069
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1474
- Math
- 1428
- Creative writing
- 1394
- Instruction following
- 1417
- Hard prompts
- 1446
- Long queries
- 1431
- Multi-turn chat
- 1423
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
25 qwen3-235b-a22b-instruct-2507 Alibaba · Apache 2.0 235B 86GB +
- Chat Elo (human votes)
- 1423(1421-1426)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 97,085
- Size
- 235B (22B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 96GB
Benchmarks by category (LMArena Elo)
- Coding
- 1473
- Math
- 1419
- Creative writing
- 1380
- Instruction following
- 1416
- Hard prompts
- 1448
- Long queries
- 1435
- Multi-turn chat
- 1438
RAM per quant (real file sizes, via unsloth)
Q2_K 86GB UD-Q2_K_XL 89GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 125GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 169GB Q6_K 193GB Q8_0 250GB -
26 deepseek-v3.2-thinking DeepSeek · MIT ? 184GB +
- Chat Elo (human votes)
- 1423(1419-1427)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1361
- Arena votes
- 41,025
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1475
- Math
- 1425
- Creative writing
- 1391
- Instruction following
- 1419
- Hard prompts
- 1446
- Long queries
- 1442
- Multi-turn chat
- 1427
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 184GB UD-IQ1_M 199GB UD-IQ2_XXS 217GB UD-IQ2_M 228GB Q2_K 245GB UD-Q2_K_XL 247GB Q3_K_M 320GB UD-Q3_K_XL 321GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 408GB Q5_K_M 476GB UD-Q5_K_XL 482GB Q6_K 551GB Q8_0 713GB -
27 deepseek-v3.2-exp DeepSeek · MIT ? ? +
- Chat Elo (human votes)
- 1423(1416-1429)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1272
- Arena votes
- 11,913
- Size
- unconfirmed
- Context
- 164K
- Multimodal
- text only
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1465
- Math
- 1417
- Creative writing
- 1410
- Instruction following
- 1416
- Hard prompts
- 1447
- Long queries
- 1442
- Multi-turn chat
- 1430
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
28 deepseek-r1-0528 DeepSeek · MIT ? 185GB +
- Chat Elo (human votes)
- 1422(1416-1428)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 18,451
- Size
- unconfirmed
- Context
- 164K
- Multimodal
- text only
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1464
- Math
- 1396
- Creative writing
- 1394
- Instruction following
- 1392
- Hard prompts
- 1433
- Long queries
- 1407
- Multi-turn chat
- 1406
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 185GB UD-IQ1_M 200GB UD-IQ2_XXS 217GB UD-IQ2_M 229GB Q2_K 245GB UD-Q2_K_XL 251GB Q3_K_M 320GB UD-Q3_K_XL 296GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 384GB Q5_K_M 476GB UD-Q5_K_XL 481GB Q6_K 551GB Q8_0 713GB -
29 kimi-k2-0905-preview Moonshot AI · Modified MIT ? ? +
- Chat Elo (human votes)
- 1418(1411-1424)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 11,770
- Size
- unconfirmed
- Context
- 262K
- Multimodal
- text only
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1468
- Math
- 1416
- Creative writing
- 1382
- Instruction following
- 1390
- Hard prompts
- 1436
- Long queries
- 1402
- Multi-turn chat
- 1404
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
30 deepseek-v3.1 DeepSeek · MIT ? 192GB +
- Chat Elo (human votes)
- 1417(1411-1423)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 14,943
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1448
- Math
- 1415
- Creative writing
- 1389
- Instruction following
- 1403
- Hard prompts
- 1433
- Long queries
- 1421
- Multi-turn chat
- 1406
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 192GB UD-IQ1_M 207GB UD-IQ2_XXS 226GB UD-IQ2_M 236GB Q2_K 246GB UD-Q2_K_XL 256GB Q3_K_M 320GB UD-Q3_K_XL 300GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 387GB Q5_K_M 476GB UD-Q5_K_XL 485GB Q6_K 551GB Q8_0 713GB -
31 kimi-k2-0711-preview Moonshot AI · Modified MIT ? ? +
- Chat Elo (human votes)
- 1417(1413-1422)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 27,608
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1460
- Math
- 1388
- Creative writing
- 1371
- Instruction following
- 1379
- Hard prompts
- 1431
- Long queries
- 1393
- Multi-turn chat
- 1420
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
32 qwen3.5-122b-a10b Alibaba · Apache 2.0 122B 34GB +
- Chat Elo (human votes)
- 1417(1413-1422)
- AA Intelligence (benchmark)
- 32.3 · code 45.7
- WebDev Elo (coding)
- 1359
- Arena votes
- 28,504
- Size
- 122B (10B active)
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 48GB
Benchmarks by category (LMArena Elo)
- Coding
- 1459
- Math
- 1424
- Creative writing
- 1367
- Instruction following
- 1407
- Hard prompts
- 1433
- Long queries
- 1419
- Multi-turn chat
- 1418
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 34GB UD-IQ2_XXS 37GB UD-IQ2_M 39GB UD-Q2_K_XL 42GB Q3_K_M 56GB UD-Q3_K_XL 57GB Q4_K_M 77GB UD-Q4_K_XL 77GB Q5_K_M 92GB UD-Q5_K_XL 92GB Q6_K 101GB Q8_0 130GB -
33 deepseek-v3.1-terminus-thinking DeepSeek · MIT ? 187GB +
- Chat Elo (human votes)
- 1417(1407-1427)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,457
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1463
- Math
- 1406
- Creative writing
- 1385
- Instruction following
- 1420
- Hard prompts
- 1445
- Long queries
- 1443
- Multi-turn chat
- 1419
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 187GB UD-IQ1_M 201GB UD-IQ2_XXS 217GB UD-IQ2_M 229GB Q2_K 246GB UD-Q2_K_XL 251GB Q3_K_M 320GB UD-Q3_K_XL 296GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 384GB Q5_K_M 476GB UD-Q5_K_XL 482GB Q6_K 551GB Q8_0 713GB -
34 deepseek-v3.1-thinking DeepSeek · MIT ? 192GB +
- Chat Elo (human votes)
- 1417(1410-1424)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 11,726
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1457
- Math
- 1414
- Creative writing
- 1404
- Instruction following
- 1419
- Hard prompts
- 1437
- Long queries
- 1446
- Multi-turn chat
- 1414
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 192GB UD-IQ1_M 207GB UD-IQ2_XXS 226GB UD-IQ2_M 236GB Q2_K 246GB UD-Q2_K_XL 256GB Q3_K_M 320GB UD-Q3_K_XL 300GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 387GB Q5_K_M 476GB UD-Q5_K_XL 485GB Q6_K 551GB Q8_0 713GB -
35 minimax-m2.7 minimax · Modified MIT ? 61GB +
- Chat Elo (human votes)
- 1417(1413-1421)
- AA Intelligence (benchmark)
- 38.1 · code 52.6
- WebDev Elo (coding)
- 1398
- Arena votes
- 49,930
- Size
- unconfirmed
- Context
- 205K
- Multimodal
- text only
- Smallest rig that fits
- 64GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1479
- Math
- 1424
- Creative writing
- 1365
- Instruction following
- 1410
- Hard prompts
- 1443
- Long queries
- 1435
- Multi-turn chat
- 1428
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 61GB UD-IQ2_XXS 65GB UD-IQ2_M 70GB UD-Q2_K_XL 75GB UD-Q3_K_XL 102GB UD-Q4_K_XL 141GB UD-Q5_K_XL 169GB Q8_0 243GB -
36 mistral-large-3 mistral · Apache 2.0 ? ? +
- Chat Elo (human votes)
- 1415(1412-1419)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1230
- Arena votes
- 50,849
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1467
- Math
- 1402
- Creative writing
- 1374
- Instruction following
- 1403
- Hard prompts
- 1433
- Long queries
- 1417
- Multi-turn chat
- 1421
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
37 deepseek-v3.1-terminus DeepSeek · MIT ? 187GB +
- Chat Elo (human votes)
- 1415(1406-1425)
- AA Intelligence (benchmark)
- 30.4 · code 43.5
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,692
- Size
- unconfirmed
- Context
- 164K
- Multimodal
- text only
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1439
- Math
- 1395
- Creative writing
- 1407
- Instruction following
- 1394
- Hard prompts
- 1424
- Long queries
- 1418
- Multi-turn chat
- 1394
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 187GB UD-IQ1_M 201GB UD-IQ2_XXS 217GB UD-IQ2_M 229GB Q2_K 246GB UD-Q2_K_XL 251GB Q3_K_M 320GB UD-Q3_K_XL 296GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 384GB Q5_K_M 476GB UD-Q5_K_XL 482GB Q6_K 551GB Q8_0 713GB -
38 qwen3-vl-235b-a22b-instruct Alibaba · Apache 2.0 235B 63GB +
- Chat Elo (human votes)
- 1415(1409-1422)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 11,499
- Size
- 235B (22B active)
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 64GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1465
- Math
- 1410
- Creative writing
- 1361
- Instruction following
- 1414
- Hard prompts
- 1440
- Long queries
- 1424
- Multi-turn chat
- 1426
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 63GB UD-IQ1_M 70GB Q2_K 86GB UD-Q2_K_XL 88GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 125GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 167GB Q6_K 193GB Q8_0 250GB -
39 hunyuan-hy3-preview Tencent · tencent-hunyuan-community 389B 160GB est +
- Chat Elo (human votes)
- 1412(1405-1420)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1357
- Arena votes
- 6,634
- Size
- 389B (52B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Tencent's Hunyuan preview. Capable MoE; preview, so expect changes. Best for: Experimentation.
Benchmarks by category (LMArena Elo)
- Coding
- 1459
- Math
- 1429
- Creative writing
- 1356
- Instruction following
- 1399
- Hard prompts
- 1438
- Long queries
- 1425
- Multi-turn chat
- 1415
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~218 GB at Q4, ~160 GB at an aggressive dynamic quant.
-
40 glm-4.5 Zhipu · MIT ? 97GB +
- Chat Elo (human votes)
- 1411(1406-1416)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 24,286
- Size
- unconfirmed
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 128GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1454
- Math
- 1413
- Creative writing
- 1373
- Instruction following
- 1405
- Hard prompts
- 1433
- Long queries
- 1416
- Multi-turn chat
- 1406
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 97GB UD-IQ1_M 108GB UD-IQ2_XXS 116GB UD-IQ2_M 122GB Q2_K 131GB UD-Q2_K_XL 135GB Q3_K_M 171GB UD-Q3_K_XL 159GB IQ4_XS 192GB Q4_K_M 216GB UD-Q4_K_XL 204GB Q5_K_M 254GB UD-Q5_K_XL 253GB Q6_K 294GB Q8_0 381GB -
41 qwen3.5-27b Alibaba · Apache 2.0 27B 9GB +
- Chat Elo (human votes)
- 1409(1404-1413)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1357
- Arena votes
- 27,314
- Size
- 27B
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1450
- Math
- 1429
- Creative writing
- 1358
- Instruction following
- 1401
- Hard prompts
- 1427
- Long queries
- 1423
- Multi-turn chat
- 1416
RAM per quant (real file sizes, via unsloth)
UD-IQ2_XXS 9GB UD-IQ2_M 10GB UD-Q2_K_XL 11GB Q3_K_M 14GB UD-Q3_K_XL 14GB IQ4_XS 15GB Q4_K_M 17GB UD-Q4_K_XL 18GB Q5_K_M 20GB UD-Q5_K_XL 20GB Q6_K 22GB Q8_0 29GB -
42 qwen3-235b-a22b-no-thinking Alibaba · Apache 2.0 235B 98GB est +
- Chat Elo (human votes)
- 1403(1399-1408)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 38,174
- Size
- 235B (22B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 128GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1446
- Math
- 1394
- Creative writing
- 1368
- Instruction following
- 1381
- Hard prompts
- 1421
- Long queries
- 1411
- Multi-turn chat
- 1411
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~133 GB at Q4, ~98 GB at an aggressive dynamic quant.
-
43 qwen3-next-80b-a3b-instruct Alibaba · Apache 2.0 80B 23GB +
- Chat Elo (human votes)
- 1401(1396-1406)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 22,859
- Size
- 80B (3B active)
- Context
- 262K
- Multimodal
- text only
- Smallest rig that fits
- 24GB GPU
Benchmarks by category (LMArena Elo)
- Coding
- 1445
- Math
- 1418
- Creative writing
- 1316
- Instruction following
- 1379
- Hard prompts
- 1421
- Long queries
- 1390
- Multi-turn chat
- 1404
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 23GB UD-IQ1_M 24GB UD-IQ2_XXS 26GB Q2_K 29GB UD-Q2_K_XL 30GB Q3_K_M 38GB UD-Q3_K_XL 36GB IQ4_XS 43GB Q4_K_M 49GB UD-Q4_K_XL 46GB Q5_K_M 57GB UD-Q5_K_XL 57GB Q6_K 66GB Q8_0 85GB -
44 longcat-flash-chat meituan · MIT ? ? +
- Chat Elo (human votes)
- 1401(1395-1408)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 11,387
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1474
- Math
- 1416
- Creative writing
- 1330
- Instruction following
- 1389
- Hard prompts
- 1426
- Long queries
- 1385
- Multi-turn chat
- 1393
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
45 qwen3-235b-a22b-thinking-2507 Alibaba · Apache 2.0 235B 86GB +
- Chat Elo (human votes)
- 1399(1393-1406)
- AA Intelligence (benchmark)
- 19.6 · code 22.1
- WebDev Elo (coding)
- n/a
- Arena votes
- 8,985
- Size
- 235B (22B active)
- Context
- 262K
- Multimodal
- text only
- Smallest rig that fits
- 96GB
Benchmarks by category (LMArena Elo)
- Coding
- 1442
- Math
- 1398
- Creative writing
- 1374
- Instruction following
- 1386
- Hard prompts
- 1418
- Long queries
- 1401
- Multi-turn chat
- 1393
RAM per quant (real file sizes, via unsloth)
Q2_K 86GB UD-Q2_K_XL 88GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 126GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 167GB Q6_K 193GB Q8_0 250GB -
46 deepseek-r1 DeepSeek · MIT ? ? +
- Chat Elo (human votes)
- 1398(1393-1403)
- AA Intelligence (benchmark)
- 18.5 · code 24.6
- WebDev Elo (coding)
- n/a
- Arena votes
- 18,524
- Size
- unconfirmed
- Context
- 164K
- Multimodal
- text only
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1445
- Math
- 1412
- Creative writing
- 1374
- Instruction following
- 1397
- Hard prompts
- 1419
- Long queries
- 1399
- Multi-turn chat
- 1410
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
47 qwen3.5-35b-a3b Alibaba · Apache 2.0 35B 11GB +
- Chat Elo (human votes)
- 1396(1391-1400)
- AA Intelligence (benchmark)
- 24.0 · code 37.0
- WebDev Elo (coding)
- 1251
- Arena votes
- 29,158
- Size
- 35B (3B active)
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1435
- Math
- 1399
- Creative writing
- 1344
- Instruction following
- 1388
- Hard prompts
- 1413
- Long queries
- 1402
- Multi-turn chat
- 1394
RAM per quant (real file sizes, via unsloth)
UD-IQ2_XXS 11GB UD-IQ2_M 11GB UD-Q2_K_XL 12GB Q3_K_M 16GB UD-Q3_K_XL 17GB Q4_K_M 22GB UD-Q4_K_XL 22GB Q5_K_M 26GB UD-Q5_K_XL 26GB Q6_K 29GB Q8_0 37GB -
48 deepseek-v3-0324 DeepSeek · MIT ? 186GB +
- Chat Elo (human votes)
- 1396(1392-1399)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 45,480
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1429
- Math
- 1370
- Creative writing
- 1390
- Instruction following
- 1379
- Hard prompts
- 1409
- Long queries
- 1393
- Multi-turn chat
- 1410
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 186GB UD-IQ1_M 196GB UD-IQ2_XXS 219GB Q2_K 244GB UD-Q2_K_XL 248GB Q3_K_M 319GB UD-Q3_K_XL 321GB Q4_K_M 404GB UD-Q4_K_XL 405GB Q5_K_M 475GB Q6_K 551GB Q8_0 713GB -
49 qwen3-vl-235b-a22b-thinking Alibaba · Apache 2.0 235B 63GB +
- Chat Elo (human votes)
- 1395(1388-1402)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 7,941
- Size
- 235B (22B active)
- Context
- 131K
- Multimodal
- yes (images)
- Smallest rig that fits
- 64GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1455
- Math
- 1405
- Creative writing
- 1338
- Instruction following
- 1383
- Hard prompts
- 1418
- Long queries
- 1406
- Multi-turn chat
- 1386
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 63GB UD-IQ1_M 70GB Q2_K 86GB UD-Q2_K_XL 88GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 125GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 167GB Q6_K 193GB Q8_0 250GB -
50 step-3.5-flash stepfun · Apache 2.0 ? ? +
- Chat Elo (human votes)
- 1395(1391-1399)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 55,991
- Size
- unconfirmed
- Context
- 262K
- Multimodal
- text only
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1450
- Math
- 1408
- Creative writing
- 1346
- Instruction following
- 1387
- Hard prompts
- 1413
- Long queries
- 1406
- Multi-turn chat
- 1397
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
51 mimo-v2-flash (non-thinking) Xiaomi · MIT ? ? +
- Chat Elo (human votes)
- 1393(1389-1396)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1330
- Arena votes
- 46,549
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1446
- Math
- 1378
- Creative writing
- 1357
- Instruction following
- 1384
- Hard prompts
- 1414
- Long queries
- 1405
- Multi-turn chat
- 1389
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
52 minimax-m2.5 minimax · Modified MIT ? 63GB +
- Chat Elo (human votes)
- 1390(1386-1394)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1386
- Arena votes
- 41,088
- Size
- unconfirmed
- Context
- 205K
- Multimodal
- text only
- Smallest rig that fits
- 64GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1444
- Math
- 1396
- Creative writing
- 1358
- Instruction following
- 1381
- Hard prompts
- 1415
- Long queries
- 1403
- Multi-turn chat
- 1394
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 63GB UD-IQ1_M 68GB UD-IQ2_XXS 74GB UD-IQ2_M 78GB Q2_K 83GB UD-Q2_K_XL 86GB Q3_K_M 109GB UD-Q3_K_XL 101GB IQ4_XS 122GB Q4_K_M 138GB UD-Q4_K_XL 131GB Q5_K_M 162GB UD-Q5_K_XL 162GB Q6_K 188GB Q8_0 243GB -
53 qwen3-coder-480b-a35b-instruct Alibaba · Apache 2.0 480B 150GB +
- Chat Elo (human votes)
- 1388(1383-1392)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1272
- Arena votes
- 25,694
- Size
- 480B (35B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1457
- Math
- 1377
- Creative writing
- 1366
- Instruction following
- 1385
- Hard prompts
- 1414
- Long queries
- 1409
- Multi-turn chat
- 1399
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 150GB Q2_K 175GB UD-Q2_K_XL 180GB Q3_K_M 229GB UD-Q3_K_XL 213GB IQ4_XS 261GB Q4_K_M 290GB UD-Q4_K_XL 276GB Q5_K_M 340GB UD-Q5_K_XL 340GB Q6_K 394GB Q8_0 510GB -
54 mimo-v2-flash (thinking) Xiaomi · MIT ? ? +
- Chat Elo (human votes)
- 1387(1381-1393)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1291
- Arena votes
- 10,934
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1431
- Math
- 1373
- Creative writing
- 1336
- Instruction following
- 1379
- Hard prompts
- 1412
- Long queries
- 1402
- Multi-turn chat
- 1375
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
55 minimax-m2.1-preview minimax · MIT ? 63GB +
- Chat Elo (human votes)
- 1384(1379-1389)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1388
- Arena votes
- 17,083
- Size
- unconfirmed
- Context
- 205K
- Multimodal
- text only
- Smallest rig that fits
- 64GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1439
- Math
- 1392
- Creative writing
- 1345
- Instruction following
- 1385
- Hard prompts
- 1407
- Long queries
- 1410
- Multi-turn chat
- 1391
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 63GB UD-IQ1_M 68GB UD-IQ2_XXS 74GB UD-IQ2_M 78GB Q2_K 83GB UD-Q2_K_XL 86GB Q3_K_M 109GB UD-Q3_K_XL 101GB IQ4_XS 122GB Q4_K_M 138GB UD-Q4_K_XL 131GB Q5_K_M 162GB UD-Q5_K_XL 162GB Q6_K 188GB Q8_0 243GB -
56 qwen3-30b-a3b-instruct-2507 Alibaba · Apache 2.0 30B 9GB +
- Chat Elo (human votes)
- 1383(1378-1388)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 23,718
- Size
- 30B (3B active)
- Context
- 262K
- Multimodal
- text only
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1439
- Math
- 1380
- Creative writing
- 1321
- Instruction following
- 1367
- Hard prompts
- 1407
- Long queries
- 1381
- Multi-turn chat
- 1383
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 9GB UD-IQ1_M 10GB UD-IQ2_XXS 10GB UD-IQ2_M 11GB Q2_K 11GB UD-Q2_K_XL 12GB Q3_K_M 15GB UD-Q3_K_XL 14GB IQ4_XS 16GB Q4_K_M 19GB UD-Q4_K_XL 18GB Q5_K_M 22GB UD-Q5_K_XL 22GB Q6_K 25GB Q8_0 32GB -
57 glm-4.6v Zhipu · MIT ? 37GB +
- Chat Elo (human votes)
- 1377(1366-1389)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 2,797
- Size
- unconfirmed
- Context
- 131K
- Multimodal
- yes (images)
- Smallest rig that fits
- 48GB
Benchmarks by category (LMArena Elo)
- Coding
- 1418
- Creative writing
- 1344
- Instruction following
- 1368
- Hard prompts
- 1383
- Long queries
- 1378
- Multi-turn chat
- 1367
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 37GB UD-IQ1_M 39GB UD-IQ2_XXS 41GB UD-IQ2_M 43GB Q2_K 44GB UD-Q2_K_XL 46GB Q3_K_M 55GB UD-Q3_K_XL 53GB IQ4_XS 58GB Q4_K_M 71GB UD-Q4_K_XL 65GB Q5_K_M 81GB UD-Q5_K_XL 80GB Q6_K 96GB Q8_0 114GB -
58 qwen3-235b-a22b Alibaba · Apache 2.0 235B 86GB +
- Chat Elo (human votes)
- 1375(1370-1380)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 26,254
- Size
- 235B (22B active)
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 96GB
Benchmarks by category (LMArena Elo)
- Coding
- 1433
- Math
- 1393
- Creative writing
- 1323
- Instruction following
- 1357
- Hard prompts
- 1392
- Long queries
- 1378
- Multi-turn chat
- 1372
RAM per quant (real file sizes, via unsloth)
Q2_K 86GB UD-Q2_K_XL 88GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 126GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 167GB Q6_K 193GB Q8_0 250GB -
59 glm-4.5-air Zhipu · MIT ? 39GB +
- Chat Elo (human votes)
- 1373(1369-1377)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 31,062
- Size
- unconfirmed
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 48GB
Benchmarks by category (LMArena Elo)
- Coding
- 1426
- Math
- 1390
- Creative writing
- 1328
- Instruction following
- 1362
- Hard prompts
- 1391
- Long queries
- 1379
- Multi-turn chat
- 1369
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 39GB UD-IQ1_M 40GB UD-IQ2_XXS 43GB UD-IQ2_M 44GB Q2_K 45GB UD-Q2_K_XL 47GB Q3_K_M 57GB UD-Q3_K_XL 55GB IQ4_XS 60GB Q4_K_M 73GB UD-Q4_K_XL 68GB Q5_K_M 84GB UD-Q5_K_XL 83GB Q6_K 99GB Q8_0 117GB -
60 qwen3-next-80b-a3b-thinking Alibaba · Apache 2.0 80B 23GB +
- Chat Elo (human votes)
- 1370(1364-1375)
- AA Intelligence (benchmark)
- 16.7 · code 17.4
- WebDev Elo (coding)
- n/a
- Arena votes
- 13,678
- Size
- 80B (3B active)
- Context
- 262K
- Multimodal
- text only
- Smallest rig that fits
- 24GB GPU
Benchmarks by category (LMArena Elo)
- Coding
- 1421
- Math
- 1389
- Creative writing
- 1324
- Instruction following
- 1359
- Hard prompts
- 1384
- Long queries
- 1370
- Multi-turn chat
- 1350
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 23GB UD-IQ1_M 24GB UD-IQ2_XXS 26GB Q2_K 29GB UD-Q2_K_XL 30GB Q3_K_M 38GB UD-Q3_K_XL 35GB IQ4_XS 43GB Q4_K_M 49GB UD-Q4_K_XL 46GB Q5_K_M 57GB UD-Q5_K_XL 57GB Q6_K 66GB Q8_0 85GB -
61 glm-4.7-flash Zhipu · MIT ? 9GB +
- Chat Elo (human votes)
- 1368(1362-1374)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 11,719
- Size
- unconfirmed
- Context
- 203K
- Multimodal
- text only
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1423
- Math
- 1366
- Creative writing
- 1313
- Instruction following
- 1351
- Hard prompts
- 1387
- Long queries
- 1374
- Multi-turn chat
- 1363
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 9GB UD-IQ1_M 10GB UD-IQ2_XXS 11GB UD-IQ2_M 11GB Q2_K 11GB UD-Q2_K_XL 12GB Q3_K_M 15GB UD-Q3_K_XL 14GB IQ4_XS 16GB Q4_K_M 18GB UD-Q4_K_XL 18GB Q5_K_M 21GB UD-Q5_K_XL 22GB Q6_K 25GB Q8_0 32GB -
62 gemma-3-27b-it Google · Gemma 27B 7GB +
- Chat Elo (human votes)
- 1366(1362-1369)
- AA Intelligence (benchmark)
- 7.4 · code 10.1
- WebDev Elo (coding)
- n/a
- Arena votes
- 47,508
- Size
- 27B
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1358
- Math
- 1322
- Creative writing
- 1348
- Instruction following
- 1344
- Hard prompts
- 1365
- Long queries
- 1364
- Multi-turn chat
- 1358
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 7GB UD-IQ1_M 7GB UD-IQ2_XXS 8GB UD-IQ2_M 10GB Q2_K 11GB UD-Q2_K_XL 11GB Q3_K_M 13GB UD-Q3_K_XL 14GB IQ4_XS 15GB Q4_K_M 17GB UD-Q4_K_XL 17GB Q5_K_M 19GB UD-Q5_K_XL 19GB Q6_K 22GB Q8_0 29GB -
63 minimax-m1 minimax · Apache 2.0 ? ? +
- Chat Elo (human votes)
- 1364(1359-1368)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 35,163
- Size
- unconfirmed
- Context
- 1M
- Multimodal
- text only
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1416
- Math
- 1371
- Creative writing
- 1318
- Instruction following
- 1346
- Hard prompts
- 1381
- Long queries
- 1365
- Multi-turn chat
- 1357
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
64 nvidia-nemotron-3-super-120b-a12b NVIDIA · NVIDIA Open Model 120B 53GB +
- Chat Elo (human votes)
- 1362(1354-1369)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 7,544
- Size
- 120B (12B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 64GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1410
- Math
- 1376
- Creative writing
- 1304
- Instruction following
- 1346
- Hard prompts
- 1381
- Long queries
- 1361
- Multi-turn chat
- 1349
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 53GB UD-IQ2_XXS 53GB UD-IQ2_M 53GB UD-Q2_K_XL 55GB UD-Q3_K_XL 63GB UD-Q4_K_XL 84GB UD-Q5_K_XL 108GB Q8_0 128GB -
65 deepseek-v3 DeepSeek · DeepSeek ? ? +
- Chat Elo (human votes)
- 1359(1354-1363)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 21,770
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1388
- Math
- 1311
- Creative writing
- 1349
- Instruction following
- 1344
- Hard prompts
- 1350
- Long queries
- 1375
- Multi-turn chat
- 1374
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
66 mistral-small-2506 mistral · Apache 2.0 ? ? +
- Chat Elo (human votes)
- 1358(1352-1363)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 17,698
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1412
- Math
- 1339
- Creative writing
- 1324
- Instruction following
- 1339
- Hard prompts
- 1374
- Long queries
- 1360
- Multi-turn chat
- 1367
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
67 command-a-03-2025 Cohere · CC-BY-NC-4.0 ? ? +
- Chat Elo (human votes)
- 1354(1350-1357)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 56,217
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1390
- Math
- 1309
- Creative writing
- 1336
- Instruction following
- 1342
- Hard prompts
- 1368
- Long queries
- 1369
- Multi-turn chat
- 1360
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
68 glm-4.5v Zhipu · MIT ? 70GB +
- Chat Elo (human votes)
- 1353(1345-1362)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 4,957
- Size
- unconfirmed
- Context
- 66K
- Multimodal
- yes (images)
- Smallest rig that fits
- 96GB
Benchmarks by category (LMArena Elo)
- Coding
- 1404
- Math
- 1358
- Creative writing
- 1310
- Instruction following
- 1342
- Hard prompts
- 1376
- Long queries
- 1341
- Multi-turn chat
- 1358
RAM per quant (real file sizes, via ggml-org)
Q4_K_M 70GB -
69 gpt-oss-120b OpenAI · Apache 2.0 120B 63GB +
- Chat Elo (human votes)
- 1352(1348-1357)
- AA Intelligence (benchmark)
- 23.8 · code 30.4
- WebDev Elo (coding)
- n/a
- Arena votes
- 30,615
- Size
- 120B
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 64GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1390
- Math
- 1381
- Creative writing
- 1278
- Instruction following
- 1325
- Hard prompts
- 1362
- Long queries
- 1324
- Multi-turn chat
- 1328
RAM per quant (real file sizes, via unsloth)
Q2_K 63GB Q3_K_M 63GB Q4_K_M 63GB UD-Q4_K_XL 63GB Q5_K_M 63GB Q6_K 63GB Q8_0 63GB -
70 step-3 stepfun · Apache 2.0 ? ? +
- Chat Elo (human votes)
- 1348(1341-1356)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 6,532
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1408
- Math
- 1363
- Creative writing
- 1308
- Instruction following
- 1345
- Hard prompts
- 1378
- Long queries
- 1348
- Multi-turn chat
- 1342
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
71 llama-3.1-nemotron-ultra-253b-v1 NVIDIA · Nvidia Open Model 253B 105GB est +
- Chat Elo (human votes)
- 1347(1336-1359)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 2,549
- Size
- 253B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 128GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1391
- Math
- 1380
- Creative writing
- 1332
- Instruction following
- 1351
- Hard prompts
- 1375
- Long queries
- 1341
- Multi-turn chat
- 1342
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~143 GB at Q4, ~105 GB at an aggressive dynamic quant.
-
72 qwen3-32b Alibaba · Apache 2.0 32B 8GB +
- Chat Elo (human votes)
- 1347(1338-1357)
- AA Intelligence (benchmark)
- 11.5 · code 15.3
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,926
- Size
- 32B
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1407
- Math
- 1399
- Creative writing
- 1305
- Instruction following
- 1332
- Hard prompts
- 1367
- Long queries
- 1355
- Multi-turn chat
- 1338
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 8GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 12GB Q2_K 12GB UD-Q2_K_XL 13GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 18GB Q4_K_M 20GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 27GB Q8_0 35GB -
73 ling-flash-2.0 ant-group · MIT ? 36GB +
- Chat Elo (human votes)
- 1346(1339-1354)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 6,997
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 48GB
Benchmarks by category (LMArena Elo)
- Coding
- 1412
- Math
- 1354
- Creative writing
- 1269
- Instruction following
- 1317
- Hard prompts
- 1366
- Long queries
- 1326
- Multi-turn chat
- 1317
RAM per quant (real file sizes, via bartowski)
Q2_K 36GB Q3_K_M 47GB -
74 minimax-m2 minimax · Apache 2.0 ? 64GB +
- Chat Elo (human votes)
- 1346(1338-1354)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- 1296
- Arena votes
- 6,862
- Size
- unconfirmed
- Context
- 205K
- Multimodal
- text only
- Smallest rig that fits
- 64GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1385
- Math
- 1355
- Creative writing
- 1287
- Instruction following
- 1339
- Hard prompts
- 1369
- Long queries
- 1343
- Multi-turn chat
- 1364
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 64GB UD-IQ1_M 69GB UD-IQ2_XXS 74GB UD-IQ2_M 78GB Q2_K 83GB UD-Q2_K_XL 86GB Q3_K_M 109GB UD-Q3_K_XL 101GB IQ4_XS 122GB Q4_K_M 138GB UD-Q4_K_XL 132GB Q5_K_M 162GB UD-Q5_K_XL 162GB Q6_K 188GB Q8_0 243GB -
75 nvidia-llama-3.3-nemotron-super-49b-v1.5 NVIDIA · Nvidia Open 49B 24GB est +
- Chat Elo (human votes)
- 1343(1333-1353)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,344
- Size
- 49B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 24GB GPU
Benchmarks by category (LMArena Elo)
- Coding
- 1405
- Math
- 1390
- Creative writing
- 1308
- Instruction following
- 1322
- Hard prompts
- 1361
- Long queries
- 1343
- Multi-turn chat
- 1342
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~31 GB at Q4, ~24 GB at an aggressive dynamic quant.
-
76 gemma-3-12b-it Google · Gemma 12B 3GB +
- Chat Elo (human votes)
- 1342(1332-1351)
- AA Intelligence (benchmark)
- 5.5 · code 5.8
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,829
- Size
- 12B
- Context
- 131K
- Multimodal
- yes (images)
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1316
- Math
- 1318
- Creative writing
- 1334
- Instruction following
- 1321
- Hard prompts
- 1332
- Long queries
- 1346
- Multi-turn chat
- 1344
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 3GB UD-IQ1_M 3GB UD-IQ2_XXS 4GB UD-IQ2_M 4GB Q2_K 5GB UD-Q2_K_XL 5GB Q3_K_M 6GB UD-Q3_K_XL 6GB IQ4_XS 7GB Q4_K_M 7GB UD-Q4_K_XL 7GB Q5_K_M 8GB UD-Q5_K_XL 8GB Q6_K 10GB Q8_0 13GB -
77 qwq-32b Alibaba · Apache 2.0 32B 8GB +
- Chat Elo (human votes)
- 1336(1332-1341)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 25,379
- Size
- 32B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1384
- Math
- 1364
- Creative writing
- 1294
- Instruction following
- 1324
- Hard prompts
- 1357
- Long queries
- 1336
- Multi-turn chat
- 1322
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 8GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 11GB Q2_K 12GB UD-Q2_K_XL 13GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 18GB Q4_K_M 20GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 27GB Q8_0 35GB -
78 llama-3.1-405b-instruct-bf16 Meta · Llama 3.1 Community 405B 166GB est +
- Chat Elo (human votes)
- 1335(1331-1339)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 41,375
- Size
- 405B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1375
- Math
- 1316
- Creative writing
- 1301
- Instruction following
- 1313
- Hard prompts
- 1341
- Long queries
- 1327
- Multi-turn chat
- 1339
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~227 GB at Q4, ~166 GB at an aggressive dynamic quant.
-
79 llama-3.1-405b-instruct-fp8 Meta · Llama 3.1 Community 405B 166GB est +
- Chat Elo (human votes)
- 1333(1329-1337)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 59,656
- Size
- 405B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1368
- Math
- 1319
- Creative writing
- 1304
- Instruction following
- 1314
- Hard prompts
- 1336
- Long queries
- 1319
- Multi-turn chat
- 1329
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~227 GB at Q4, ~166 GB at an aggressive dynamic quant.
-
80 olmo-3.1-32b-instruct Allen AI · Apache 2.0 32B 7GB +
- Chat Elo (human votes)
- 1330(1324-1336)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 12,211
- Size
- 32B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1384
- Math
- 1305
- Creative writing
- 1291
- Instruction following
- 1322
- Hard prompts
- 1350
- Long queries
- 1337
- Multi-turn chat
- 1328
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 7GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 11GB Q2_K 12GB UD-Q2_K_XL 12GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 17GB Q4_K_M 19GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 26GB Q8_0 34GB -
81 molmo-2-8b Allen AI · Apache 2.0 8B 7GB est +
- Chat Elo (human votes)
- 1328(1307-1350)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 799
- Size
- 8B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Instruction following
- 1313
- Hard prompts
- 1341
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
82 llama-3.3-nemotron-49b-super-v1 NVIDIA · Nvidia 49B 24GB est +
- Chat Elo (human votes)
- 1328(1316-1340)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 2,218
- Size
- 49B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 24GB GPU
Benchmarks by category (LMArena Elo)
- Coding
- 1363
- Creative writing
- 1298
- Instruction following
- 1326
- Hard prompts
- 1361
- Long queries
- 1338
- Multi-turn chat
- 1331
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~31 GB at Q4, ~24 GB at an aggressive dynamic quant.
-
83 qwen3-30b-a3b Alibaba · Apache 2.0 30B 9GB +
- Chat Elo (human votes)
- 1327(1322-1332)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 26,470
- Size
- 30B (3B active)
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1387
- Math
- 1352
- Creative writing
- 1284
- Instruction following
- 1312
- Hard prompts
- 1345
- Long queries
- 1339
- Multi-turn chat
- 1320
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 9GB UD-IQ1_M 10GB UD-IQ2_XXS 10GB UD-IQ2_M 11GB Q2_K 11GB UD-Q2_K_XL 12GB Q3_K_M 15GB UD-Q3_K_XL 14GB IQ4_XS 16GB Q4_K_M 19GB UD-Q4_K_XL 18GB Q5_K_M 22GB UD-Q5_K_XL 22GB Q6_K 25GB Q8_0 32GB -
84 llama-4-maverick-17b-128e-instruct Meta · Llama 4 17B 121GB +
- Chat Elo (human votes)
- 1327(1323-1331)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 39,950
- Size
- 17B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 128GB Mac
Benchmarks by category (LMArena Elo)
- Coding
- 1373
- Math
- 1318
- Creative writing
- 1307
- Instruction following
- 1314
- Hard prompts
- 1339
- Long queries
- 1335
- Multi-turn chat
- 1324
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 121GB UD-IQ1_M 128GB UD-IQ2_XXS 135GB UD-IQ2_M 141GB Q2_K 146GB UD-Q2_K_XL 153GB Q3_K_M 191GB UD-Q3_K_XL 180GB IQ4_XS 214GB Q4_K_M 243GB UD-Q4_K_XL 232GB Q5_K_M 284GB UD-Q5_K_XL 287GB Q6_K 329GB Q8_0 426GB -
85 deepseek-v2.5-1210 DeepSeek · DeepSeek ? 142GB +
- Chat Elo (human votes)
- 1323(1315-1332)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 6,795
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1375
- Math
- 1292
- Creative writing
- 1310
- Instruction following
- 1315
- Hard prompts
- 1329
- Long queries
- 1337
- Multi-turn chat
- 1323
RAM per quant (real file sizes, via lmstudio-community)
Q4_K_M 142GB Q6_K 194GB Q8_0 251GB -
86 llama-4-scout-17b-16e-instruct Meta · Llama 17B 32GB +
- Chat Elo (human votes)
- 1323(1318-1327)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 30,265
- Size
- 17B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1362
- Math
- 1309
- Creative writing
- 1290
- Instruction following
- 1301
- Hard prompts
- 1330
- Long queries
- 1327
- Multi-turn chat
- 1321
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 32GB UD-IQ1_M 35GB UD-IQ2_XXS 37GB UD-IQ2_M 39GB Q2_K 40GB UD-Q2_K_XL 42GB Q3_K_M 52GB UD-Q3_K_XL 49GB IQ4_XS 58GB Q4_K_M 65GB UD-Q4_K_XL 62GB Q5_K_M 77GB UD-Q5_K_XL 79GB Q6_K 88GB Q8_0 115GB -
87 ring-flash-2.0 ant-group · MIT ? ? +
- Chat Elo (human votes)
- 1321(1313-1328)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 7,138
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1391
- Math
- 1339
- Creative writing
- 1260
- Instruction following
- 1317
- Hard prompts
- 1350
- Long queries
- 1330
- Multi-turn chat
- 1283
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
88 llama-3.3-70b-instruct Meta · Llama-3.3 70B 16GB +
- Chat Elo (human votes)
- 1318(1315-1322)
- AA Intelligence (benchmark)
- 9.4 · code 11.9
- WebDev Elo (coding)
- n/a
- Arena votes
- 54,723
- Size
- 70B
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1345
- Math
- 1296
- Creative writing
- 1286
- Instruction following
- 1292
- Hard prompts
- 1321
- Long queries
- 1312
- Multi-turn chat
- 1317
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 16GB UD-IQ1_M 17GB UD-IQ2_XXS 19GB UD-IQ2_M 24GB Q2_K 26GB UD-Q2_K_XL 27GB Q3_K_M 34GB UD-Q3_K_XL 35GB IQ4_XS 38GB Q4_K_M 43GB UD-Q4_K_XL 43GB Q5_K_M 50GB UD-Q5_K_XL 50GB Q6_K 58GB -
89 gemma-3n-e4b-it Google · Gemma 4B 3GB +
- Chat Elo (human votes)
- 1318(1313-1323)
- AA Intelligence (benchmark)
- n/a · code 3.2
- WebDev Elo (coding)
- n/a
- Arena votes
- 22,565
- Size
- 4B
- Context
- 33K
- Multimodal
- text only
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1308
- Math
- 1260
- Creative writing
- 1300
- Instruction following
- 1281
- Hard prompts
- 1313
- Long queries
- 1311
- Multi-turn chat
- 1292
RAM per quant (real file sizes, via unsloth)
UD-IQ2_XXS 3GB UD-IQ2_M 3GB Q2_K 3GB UD-Q2_K_XL 4GB Q3_K_M 4GB UD-Q3_K_XL 4GB IQ4_XS 4GB Q4_K_M 5GB UD-Q4_K_XL 5GB Q5_K_M 5GB UD-Q5_K_XL 6GB Q6_K 6GB Q8_0 7GB -
90 qwen-max-0919 Alibaba · Qwen ? ? +
- Chat Elo (human votes)
- 1318(1312-1324)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 16,478
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1353
- Math
- 1291
- Creative writing
- 1285
- Instruction following
- 1302
- Hard prompts
- 1319
- Long queries
- 1326
- Multi-turn chat
- 1305
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
91 gpt-oss-20b OpenAI · Apache 2.0 20B 11GB +
- Chat Elo (human votes)
- 1317(1311-1324)
- AA Intelligence (benchmark)
- 14.9 · code 20.7
- WebDev Elo (coding)
- n/a
- Arena votes
- 10,621
- Size
- 20B
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1369
- Math
- 1336
- Creative writing
- 1238
- Instruction following
- 1281
- Hard prompts
- 1323
- Long queries
- 1301
- Multi-turn chat
- 1291
RAM per quant (real file sizes, via unsloth)
Q2_K 11GB Q3_K_M 12GB Q4_K_M 12GB UD-Q4_K_XL 12GB Q5_K_M 12GB Q6_K 12GB Q8_0 12GB -
92 nvidia-nemotron-3-nano-30b-a3b-bf16 NVIDIA · NVIDIA Open Model 30B 16GB est +
- Chat Elo (human votes)
- 1316(1310-1321)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 15,494
- Size
- 30B (3B active)
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1363
- Math
- 1352
- Creative writing
- 1249
- Instruction following
- 1293
- Hard prompts
- 1328
- Long queries
- 1291
- Multi-turn chat
- 1299
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~21 GB at Q4, ~16 GB at an aggressive dynamic quant.
-
93 mistral-large-2407 mistral · Mistral Research ? ? +
- Chat Elo (human votes)
- 1314(1310-1318)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 45,459
- Size
- unconfirmed
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1354
- Math
- 1288
- Creative writing
- 1287
- Instruction following
- 1299
- Hard prompts
- 1320
- Long queries
- 1304
- Multi-turn chat
- 1296
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
94 deepseek-v2.5 DeepSeek · DeepSeek ? 142GB +
- Chat Elo (human votes)
- 1307(1302-1312)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 24,572
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1368
- Math
- 1288
- Creative writing
- 1265
- Instruction following
- 1292
- Hard prompts
- 1322
- Long queries
- 1314
- Multi-turn chat
- 1292
RAM per quant (real file sizes, via lmstudio-community)
Q4_K_M 142GB Q6_K 194GB Q8_0 251GB -
95 granite-4.1-8b IBM · Apache 2.0 8B 3GB +
- Chat Elo (human votes)
- 1307(1297-1317)
- AA Intelligence (benchmark)
- n/a · code 9.5
- WebDev Elo (coding)
- 1194
- Arena votes
- 4,062
- Size
- 8B
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1353
- Math
- 1321
- Creative writing
- 1268
- Instruction following
- 1290
- Hard prompts
- 1323
- Long queries
- 1296
- Multi-turn chat
- 1286
RAM per quant (real file sizes, via unsloth)
UD-IQ2_M 3GB UD-Q2_K_XL 4GB Q3_K_M 4GB UD-Q3_K_XL 5GB IQ4_XS 5GB Q4_K_M 5GB UD-Q4_K_XL 5GB Q5_K_M 6GB UD-Q5_K_XL 6GB Q6_K 7GB Q8_0 9GB -
96 olmo-3-32b-think Allen AI · Apache 2.0 32B 7GB +
- Chat Elo (human votes)
- 1306(1297-1314)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 5,933
- Size
- 32B
- Context
- 66K
- Multimodal
- text only
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1364
- Math
- 1311
- Creative writing
- 1262
- Instruction following
- 1300
- Hard prompts
- 1328
- Long queries
- 1320
- Multi-turn chat
- 1300
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 7GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 11GB Q2_K 12GB UD-Q2_K_XL 12GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 17GB Q4_K_M 19GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 26GB Q8_0 34GB -
97 mistral-large-2411 mistral · MRL ? ? +
- Chat Elo (human votes)
- 1305(1301-1310)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 28,073
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1346
- Math
- 1282
- Creative writing
- 1276
- Instruction following
- 1295
- Hard prompts
- 1313
- Long queries
- 1305
- Multi-turn chat
- 1293
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
98 gemma-3-4b-it Google · Gemma 4B 1GB +
- Chat Elo (human votes)
- 1303(1294-1313)
- AA Intelligence (benchmark)
- n/a · code 2.7
- WebDev Elo (coding)
- n/a
- Arena votes
- 4,171
- Size
- 4B
- Context
- 131K
- Multimodal
- yes (images)
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1274
- Math
- 1254
- Creative writing
- 1276
- Instruction following
- 1268
- Hard prompts
- 1284
- Long queries
- 1311
- Multi-turn chat
- 1272
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 1GB UD-IQ1_M 1GB UD-IQ2_XXS 1GB UD-IQ2_M 2GB Q2_K 2GB UD-Q2_K_XL 2GB Q3_K_M 2GB UD-Q3_K_XL 2GB IQ4_XS 2GB Q4_K_M 2GB UD-Q4_K_XL 3GB Q5_K_M 3GB UD-Q5_K_XL 3GB Q6_K 3GB Q8_0 4GB -
99 mistral-small-3.1-24b-instruct-2503 mistral · Apache 2.0 24B 6GB +
- Chat Elo (human votes)
- 1303(1299-1308)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 33,189
- Size
- 24B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1362
- Math
- 1278
- Creative writing
- 1271
- Instruction following
- 1295
- Hard prompts
- 1319
- Long queries
- 1322
- Multi-turn chat
- 1291
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 6GB UD-IQ1_M 6GB UD-IQ2_XXS 7GB UD-IQ2_M 8GB Q2_K 9GB UD-Q2_K_XL 9GB Q3_K_M 11GB UD-Q3_K_XL 12GB IQ4_XS 13GB Q4_K_M 14GB UD-Q4_K_XL 15GB Q5_K_M 17GB UD-Q5_K_XL 17GB Q6_K 19GB Q8_0 25GB -
100 qwen2.5-72b-instruct Alibaba · Qwen 72B 47GB +
- Chat Elo (human votes)
- 1303(1299-1307)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 39,406
- Size
- 72B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 48GB
Benchmarks by category (LMArena Elo)
- Coding
- 1356
- Math
- 1296
- Creative writing
- 1254
- Instruction following
- 1292
- Hard prompts
- 1318
- Long queries
- 1317
- Multi-turn chat
- 1299
RAM per quant (real file sizes, via lmstudio-community)
Q4_K_M 47GB Q6_K 64GB Q8_0 77GB -
101 llama-3.1-nemotron-70b-instruct NVIDIA · Llama 3.1 70B 32GB est +
- Chat Elo (human votes)
- 1299(1291-1307)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 7,140
- Size
- 70B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1328
- Math
- 1279
- Creative writing
- 1277
- Instruction following
- 1280
- Hard prompts
- 1310
- Long queries
- 1281
- Multi-turn chat
- 1293
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.
-
102 llama-3.1-70b-instruct Meta · Llama 3.1 Community 70B 32GB est +
- Chat Elo (human votes)
- 1293(1289-1297)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 55,240
- Size
- 70B
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1333
- Math
- 1269
- Creative writing
- 1257
- Instruction following
- 1272
- Hard prompts
- 1298
- Long queries
- 1294
- Multi-turn chat
- 1288
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.
-
103 gemma-2-27b-it Google · Gemma license 27B 15GB +
- Chat Elo (human votes)
- 1289(1286-1292)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 75,754
- Size
- 27B
- Context
- 8K
- Multimodal
- text only
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1305
- Math
- 1246
- Creative writing
- 1292
- Instruction following
- 1271
- Hard prompts
- 1282
- Long queries
- 1299
- Multi-turn chat
- 1279
RAM per quant (real file sizes, via lmstudio-community)
IQ4_XS 15GB Q4_K_M 17GB Q5_K_M 19GB Q6_K 22GB Q8_0 29GB -
104 ibm-granite-h-small IBM · Apache 2.0 ? ? +
- Chat Elo (human votes)
- 1287(1279-1296)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 5,682
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1329
- Math
- 1279
- Creative writing
- 1243
- Instruction following
- 1268
- Hard prompts
- 1301
- Long queries
- 1291
- Multi-turn chat
- 1282
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
105 llama-3.1-nemotron-51b-instruct NVIDIA · Llama 3.1 51B 24GB est +
- Chat Elo (human votes)
- 1286(1276-1296)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,749
- Size
- 51B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 24GB GPU
Benchmarks by category (LMArena Elo)
- Coding
- 1311
- Math
- 1271
- Creative writing
- 1261
- Instruction following
- 1261
- Hard prompts
- 1281
- Long queries
- 1270
- Multi-turn chat
- 1276
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~32 GB at Q4, ~24 GB at an aggressive dynamic quant.
-
106 llama-3.1-tulu-3-70b Allen AI · Llama 3.1 70B 26GB +
- Chat Elo (human votes)
- 1286(1275-1296)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 2,846
- Size
- 70B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1308
- Math
- 1264
- Creative writing
- 1247
- Instruction following
- 1272
- Hard prompts
- 1273
- Long queries
- 1272
- Multi-turn chat
- 1278
RAM per quant (real file sizes, via unsloth)
Q2_K 26GB Q3_K_M 34GB Q4_K_M 43GB Q5_K_M 50GB -
107 olmo-3.1-32b-think Allen AI · Apache 2.0 32B 7GB +
- Chat Elo (human votes)
- 1285(1278-1292)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 8,495
- Size
- 32B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1338
- Math
- 1297
- Creative writing
- 1247
- Instruction following
- 1278
- Hard prompts
- 1305
- Long queries
- 1302
- Multi-turn chat
- 1266
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 7GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 11GB Q2_K 12GB UD-Q2_K_XL 12GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 17GB Q4_K_M 19GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 26GB Q8_0 34GB -
108 nemotron-4-340b-instruct NVIDIA · NVIDIA Open Model 340B 140GB est +
- Chat Elo (human votes)
- 1277(1271-1282)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 19,659
- Size
- 340B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 192GB
Benchmarks by category (LMArena Elo)
- Coding
- 1308
- Math
- 1252
- Creative writing
- 1238
- Instruction following
- 1258
- Hard prompts
- 1280
- Long queries
- 1288
- Multi-turn chat
- 1256
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~191 GB at Q4, ~140 GB at an aggressive dynamic quant.
-
109 llama-3-70b-instruct Meta · Llama 3 Community 70B 32GB est +
- Chat Elo (human votes)
- 1276(1272-1280)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 156,876
- Size
- 70B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1306
- Math
- 1258
- Creative writing
- 1254
- Instruction following
- 1258
- Hard prompts
- 1279
- Long queries
- 1252
- Multi-turn chat
- 1272
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.
-
110 command-r-plus-08-2024 Cohere · CC-BY-NC-4.0 ? ? +
- Chat Elo (human votes)
- 1276(1269-1282)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 9,866
- Size
- unconfirmed
- Context
- 128K
- Multimodal
- text only
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1280
- Math
- 1231
- Creative writing
- 1263
- Instruction following
- 1253
- Hard prompts
- 1260
- Long queries
- 1281
- Multi-turn chat
- 1249
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
111 mistral-small-24b-instruct-2501 mistral · Apache 2.0 24B 9GB +
- Chat Elo (human votes)
- 1274(1268-1280)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 14,681
- Size
- 24B
- Context
- 33K
- Multimodal
- text only
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1312
- Math
- 1262
- Creative writing
- 1227
- Instruction following
- 1255
- Hard prompts
- 1285
- Long queries
- 1279
- Multi-turn chat
- 1249
RAM per quant (real file sizes, via unsloth)
Q2_K 9GB Q3_K_M 11GB Q4_K_M 14GB Q6_K 19GB Q8_0 25GB -
112 qwen2.5-coder-32b-instruct Alibaba · Apache 2.0 32B 12GB +
- Chat Elo (human votes)
- 1270(1262-1279)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 5,432
- Size
- 32B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1342
- Math
- 1270
- Creative writing
- 1207
- Instruction following
- 1264
- Hard prompts
- 1303
- Long queries
- 1287
- Multi-turn chat
- 1251
RAM per quant (real file sizes, via unsloth)
Q2_K 12GB Q3_K_M 16GB Q4_K_M 20GB Q5_K_M 23GB Q6_K 27GB Q8_0 35GB -
113 c4ai-aya-expanse-32b Cohere · CC-BY-NC-4.0 32B 17GB est +
- Chat Elo (human votes)
- 1267(1262-1272)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 27,124
- Size
- 32B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 24GB GPU
Benchmarks by category (LMArena Elo)
- Coding
- 1292
- Math
- 1233
- Creative writing
- 1228
- Instruction following
- 1249
- Hard prompts
- 1271
- Long queries
- 1289
- Multi-turn chat
- 1232
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~22 GB at Q4, ~17 GB at an aggressive dynamic quant.
-
114 gemma-2-9b-it Google · Gemma license 9B 5GB +
- Chat Elo (human votes)
- 1266(1263-1270)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 54,611
- Size
- 9B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1271
- Math
- 1219
- Creative writing
- 1258
- Instruction following
- 1244
- Hard prompts
- 1257
- Long queries
- 1265
- Multi-turn chat
- 1251
RAM per quant (real file sizes, via lmstudio-community)
IQ4_XS 5GB Q4_K_M 6GB Q5_K_M 7GB Q6_K 8GB Q8_0 10GB -
115 deepseek-coder-v2 DeepSeek · DeepSeek License ? ? +
- Chat Elo (human votes)
- 1265(1258-1271)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 15,147
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1342
- Math
- 1272
- Creative writing
- 1204
- Instruction following
- 1252
- Hard prompts
- 1288
- Long queries
- 1287
- Multi-turn chat
- 1235
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
116 qwen2-72b-instruct Alibaba · Qianwen LICENSE 72B 30GB +
- Chat Elo (human votes)
- 1261(1256-1266)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 37,325
- Size
- 72B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1296
- Math
- 1273
- Creative writing
- 1223
- Instruction following
- 1241
- Hard prompts
- 1273
- Long queries
- 1260
- Multi-turn chat
- 1243
RAM per quant (real file sizes, via bartowski)
Q2_K 30GB Q3_K_M 38GB IQ4_XS 40GB Q4_K_M 47GB -
117 command-r-plus Cohere · CC-BY-NC-4.0 ? ? +
- Chat Elo (human votes)
- 1261(1257-1265)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 77,554
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1272
- Math
- 1213
- Creative writing
- 1236
- Instruction following
- 1240
- Hard prompts
- 1251
- Long queries
- 1261
- Multi-turn chat
- 1236
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
118 phi-4 Microsoft · MIT ? 6GB +
- Chat Elo (human votes)
- 1256(1251-1261)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 24,126
- Size
- unconfirmed
- Context
- 16K
- Multimodal
- text only
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1306
- Math
- 1265
- Creative writing
- 1210
- Instruction following
- 1245
- Hard prompts
- 1278
- Long queries
- 1266
- Multi-turn chat
- 1241
RAM per quant (real file sizes, via unsloth)
Q2_K 6GB Q3_K_M 7GB Q4_K_M 9GB Q5_K_M 10GB Q6_K 12GB Q8_0 16GB -
119 olmo-2-0325-32b-instruct Allen AI · Apache-2.0 32B 12GB +
- Chat Elo (human votes)
- 1251(1241-1262)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,334
- Size
- 32B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1279
- Math
- 1227
- Creative writing
- 1224
- Instruction following
- 1228
- Hard prompts
- 1261
- Long queries
- 1243
- Multi-turn chat
- 1243
RAM per quant (real file sizes, via unsloth)
Q2_K 12GB Q3_K_M 16GB Q4_K_M 19GB Q5_K_M 23GB Q6_K 26GB Q8_0 34GB -
120 command-r-08-2024 Cohere · CC-BY-NC-4.0 ? ? +
- Chat Elo (human votes)
- 1250(1243-1256)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 10,140
- Size
- unconfirmed
- Context
- 128K
- Multimodal
- text only
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1281
- Math
- 1206
- Creative writing
- 1209
- Instruction following
- 1235
- Hard prompts
- 1256
- Long queries
- 1262
- Multi-turn chat
- 1211
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
121 ministral-8b-2410 mistral · MRL 8B 7GB est +
- Chat Elo (human votes)
- 1237(1228-1246)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 4,781
- Size
- 8B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1274
- Math
- 1214
- Creative writing
- 1221
- Instruction following
- 1211
- Hard prompts
- 1251
- Long queries
- 1259
- Multi-turn chat
- 1207
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
122 qwen1.5-110b-chat Alibaba · Qianwen LICENSE 110B 41GB +
- Chat Elo (human votes)
- 1234(1228-1239)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 26,195
- Size
- 110B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 48GB
Benchmarks by category (LMArena Elo)
- Coding
- 1279
- Math
- 1221
- Creative writing
- 1193
- Instruction following
- 1217
- Hard prompts
- 1245
- Long queries
- 1228
- Multi-turn chat
- 1212
RAM per quant (real file sizes, via bartowski)
Q2_K 41GB -
123 qwen1.5-72b-chat Alibaba · Qianwen LICENSE 72B 33GB est +
- Chat Elo (human votes)
- 1233(1227-1238)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 39,302
- Size
- 72B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 48GB
Benchmarks by category (LMArena Elo)
- Coding
- 1274
- Math
- 1209
- Creative writing
- 1188
- Instruction following
- 1211
- Hard prompts
- 1239
- Long queries
- 1235
- Multi-turn chat
- 1217
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~44 GB at Q4, ~33 GB at an aggressive dynamic quant.
-
124 mixtral-8x22b-instruct-v0.1 mistral · Apache 2.0 22B 13GB est +
- Chat Elo (human votes)
- 1229(1224-1233)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 51,416
- Size
- 22B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1277
- Math
- 1228
- Creative writing
- 1190
- Instruction following
- 1214
- Hard prompts
- 1243
- Long queries
- 1217
- Multi-turn chat
- 1187
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~16 GB at Q4, ~13 GB at an aggressive dynamic quant.
-
125 command-r Cohere · CC-BY-NC-4.0 ? ? +
- Chat Elo (human votes)
- 1226(1221-1231)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 54,036
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1243
- Math
- 1176
- Creative writing
- 1195
- Instruction following
- 1198
- Hard prompts
- 1214
- Long queries
- 1232
- Multi-turn chat
- 1196
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
126 llama-3-8b-instruct Meta · Llama 3 Community 8B 7GB est +
- Chat Elo (human votes)
- 1223(1219-1227)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 104,642
- Size
- 8B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1253
- Math
- 1193
- Creative writing
- 1196
- Instruction following
- 1192
- Hard prompts
- 1219
- Long queries
- 1207
- Multi-turn chat
- 1205
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
127 c4ai-aya-expanse-8b Cohere · CC-BY-NC-4.0 8B 7GB est +
- Chat Elo (human votes)
- 1223(1216-1230)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 9,818
- Size
- 8B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1255
- Math
- 1201
- Creative writing
- 1189
- Instruction following
- 1204
- Hard prompts
- 1230
- Long queries
- 1248
- Multi-turn chat
- 1199
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
128 llama-3.1-tulu-3-8b Allen AI · Llama 3.1 8B 3GB +
- Chat Elo (human votes)
- 1220(1210-1231)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 2,896
- Size
- 8B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1253
- Math
- 1206
- Creative writing
- 1197
- Instruction following
- 1208
- Hard prompts
- 1223
- Long queries
- 1226
- Multi-turn chat
- 1178
RAM per quant (real file sizes, via unsloth)
Q2_K 3GB Q3_K_M 4GB Q4_K_M 5GB Q5_K_M 6GB Q6_K 7GB Q8_0 9GB -
129 llama-3.1-8b-instruct Meta · Llama 3.1 Community 8B 2GB +
- Chat Elo (human votes)
- 1211(1207-1215)
- AA Intelligence (benchmark)
- 7.6 · code 5.4
- WebDev Elo (coding)
- n/a
- Arena votes
- 49,605
- Size
- 8B
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1260
- Math
- 1189
- Creative writing
- 1177
- Instruction following
- 1191
- Hard prompts
- 1222
- Long queries
- 1223
- Multi-turn chat
- 1198
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 2GB UD-IQ1_M 2GB UD-IQ2_XXS 3GB UD-IQ2_M 3GB Q2_K 3GB UD-Q2_K_XL 3GB Q3_K_M 4GB UD-Q3_K_XL 4GB IQ4_XS 4GB Q4_K_M 5GB UD-Q4_K_XL 5GB Q5_K_M 6GB UD-Q5_K_XL 6GB Q6_K 7GB -
130 granite-3.1-8b-instruct IBM · Apache 2.0 8B 5GB +
- Chat Elo (human votes)
- 1208(1197-1219)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,090
- Size
- 8B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1287
- Math
- 1190
- Creative writing
- 1170
- Instruction following
- 1192
- Hard prompts
- 1230
- Long queries
- 1231
- Multi-turn chat
- 1155
RAM per quant (real file sizes, via lmstudio-community)
Q4_K_M 5GB Q6_K 7GB Q8_0 9GB -
131 qwen1.5-32b-chat Alibaba · Qianwen LICENSE 32B 12GB +
- Chat Elo (human votes)
- 1203(1197-1209)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 21,741
- Size
- 32B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1261
- Math
- 1201
- Creative writing
- 1131
- Instruction following
- 1183
- Hard prompts
- 1220
- Long queries
- 1223
- Multi-turn chat
- 1192
RAM per quant (real file sizes, via lmstudio-community)
Q2_K 12GB Q3_K_M 16GB IQ4_XS 18GB Q4_K_M 20GB Q5_K_M 23GB Q6_K 27GB Q8_0 35GB -
132 gemma-2-2b-it Google · Gemma license 2B 2GB +
- Chat Elo (human votes)
- 1200(1196-1204)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 46,616
- Size
- 2B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1193
- Math
- 1163
- Creative writing
- 1183
- Instruction following
- 1171
- Hard prompts
- 1185
- Long queries
- 1186
- Multi-turn chat
- 1164
RAM per quant (real file sizes, via lmstudio-community)
IQ4_XS 2GB Q4_K_M 2GB Q5_K_M 2GB Q6_K 2GB Q8_0 3GB -
133 phi-3-medium-4k-instruct Microsoft · MIT ? 5GB +
- Chat Elo (human votes)
- 1197(1192-1202)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 25,055
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1230
- Math
- 1216
- Creative writing
- 1149
- Instruction following
- 1176
- Hard prompts
- 1212
- Long queries
- 1190
- Multi-turn chat
- 1139
RAM per quant (real file sizes, via bartowski)
Q2_K 5GB Q3_K_M 7GB IQ4_XS 7GB Q4_K_M 9GB Q5_K_M 10GB Q6_K 11GB Q8_0 15GB -
134 mixtral-8x7b-instruct-v0.1 mistral · Apache 2.0 7B 7GB est +
- Chat Elo (human votes)
- 1196(1192-1201)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 73,503
- Size
- 7B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1239
- Math
- 1191
- Creative writing
- 1159
- Instruction following
- 1179
- Hard prompts
- 1210
- Long queries
- 1181
- Multi-turn chat
- 1166
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
135 qwen1.5-14b-chat Alibaba · Qianwen LICENSE 14B 10GB est +
- Chat Elo (human votes)
- 1190(1183-1197)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 17,839
- Size
- 14B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1239
- Math
- 1167
- Creative writing
- 1137
- Instruction following
- 1167
- Hard prompts
- 1199
- Long queries
- 1191
- Multi-turn chat
- 1164
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~12 GB at Q4, ~10 GB at an aggressive dynamic quant.
-
136 wizardlm-70b Microsoft · Llama 2 Community 70B 32GB est +
- Chat Elo (human votes)
- 1184(1175-1193)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 8,214
- Size
- 70B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1193
- Math
- 1157
- Creative writing
- 1198
- Instruction following
- 1162
- Hard prompts
- 1172
- Long queries
- 1181
- Multi-turn chat
- 1165
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.
-
137 deepseek-llm-67b-chat DeepSeek · DeepSeek License 67B 31GB est +
- Chat Elo (human votes)
- 1184(1172-1195)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 4,932
- Size
- 67B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1217
- Math
- 1156
- Creative writing
- 1135
- Instruction following
- 1156
- Hard prompts
- 1174
- Long queries
- 1183
- Multi-turn chat
- 1151
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~41 GB at Q4, ~31 GB at an aggressive dynamic quant.
-
138 granite-3.0-8b-instruct IBM · Apache 2.0 8B 5GB +
- Chat Elo (human votes)
- 1182(1173-1191)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 6,638
- Size
- 8B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1240
- Math
- 1197
- Creative writing
- 1136
- Instruction following
- 1171
- Hard prompts
- 1202
- Long queries
- 1202
- Multi-turn chat
- 1136
RAM per quant (real file sizes, via lmstudio-community)
Q4_K_M 5GB Q6_K 7GB Q8_0 9GB -
139 gemma-1.1-7b-it Google · Gemma license 7B 3GB +
- Chat Elo (human votes)
- 1182(1176-1188)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 23,893
- Size
- 7B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1217
- Math
- 1161
- Creative writing
- 1149
- Instruction following
- 1155
- Hard prompts
- 1193
- Long queries
- 1166
- Multi-turn chat
- 1124
RAM per quant (real file sizes, via bartowski)
Q2_K 3GB Q3_K_M 4GB IQ4_XS 5GB Q4_K_M 5GB Q5_K_M 6GB Q6_K 7GB Q8_0 9GB -
140 granite-3.1-2b-instruct IBM · Apache 2.0 2B 2GB +
- Chat Elo (human votes)
- 1178(1167-1190)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,188
- Size
- 2B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1248
- Math
- 1198
- Creative writing
- 1147
- Instruction following
- 1172
- Hard prompts
- 1219
- Long queries
- 1222
- Multi-turn chat
- 1142
RAM per quant (real file sizes, via lmstudio-community)
Q4_K_M 2GB Q6_K 2GB Q8_0 3GB -
141 phi-3-small-8k-instruct Microsoft · MIT ? ? +
- Chat Elo (human votes)
- 1170(1164-1176)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 17,766
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1204
- Math
- 1193
- Creative writing
- 1133
- Instruction following
- 1153
- Hard prompts
- 1188
- Long queries
- 1161
- Multi-turn chat
- 1118
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
142 llama-2-70b-chat Meta · Llama 2 Community 70B 32GB est +
- Chat Elo (human votes)
- 1170(1165-1176)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 38,492
- Size
- 70B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1178
- Math
- 1137
- Creative writing
- 1111
- Instruction following
- 1134
- Hard prompts
- 1159
- Long queries
- 1137
- Multi-turn chat
- 1133
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.
-
143 llama-3.2-3b-instruct Meta · Llama 3.2 3B 1GB +
- Chat Elo (human votes)
- 1166(1159-1174)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 7,936
- Size
- 3B
- Context
- 131K
- Multimodal
- text only
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1176
- Math
- 1165
- Creative writing
- 1144
- Instruction following
- 1145
- Hard prompts
- 1167
- Long queries
- 1158
- Multi-turn chat
- 1148
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 1GB UD-IQ1_M 1GB UD-IQ2_XXS 1GB UD-IQ2_M 1GB Q2_K 1GB UD-Q2_K_XL 1GB Q3_K_M 2GB UD-Q3_K_XL 2GB IQ4_XS 2GB Q4_K_M 2GB UD-Q4_K_XL 2GB Q5_K_M 2GB UD-Q5_K_XL 2GB Q6_K 3GB Q8_0 3GB -
144 granite-3.0-2b-instruct IBM · Apache 2.0 2B 2GB +
- Chat Elo (human votes)
- 1156(1147-1164)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 6,837
- Size
- 2B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1209
- Math
- 1169
- Creative writing
- 1099
- Instruction following
- 1131
- Hard prompts
- 1176
- Long queries
- 1147
- Multi-turn chat
- 1115
RAM per quant (real file sizes, via lmstudio-community)
Q4_K_M 2GB Q6_K 2GB Q8_0 3GB -
145 qwq-32b-preview Alibaba · Apache 2.0 32B 12GB +
- Chat Elo (human votes)
- 1155(1143-1166)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,231
- Size
- 32B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1173
- Math
- 1211
- Creative writing
- 1105
- Instruction following
- 1146
- Hard prompts
- 1168
- Long queries
- 1178
- Multi-turn chat
- 1144
RAM per quant (real file sizes, via unsloth)
Q2_K 12GB Q3_K_M 16GB Q4_K_M 20GB Q5_K_M 23GB Q6_K 27GB Q8_0 35GB -
146 llama2-70b-steerlm-chat NVIDIA · Llama 2 Community 70B 32GB est +
- Chat Elo (human votes)
- 1154(1141-1167)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 3,585
- Size
- 70B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Coding
- 1145
- Math
- 1114
- Creative writing
- 1119
- Instruction following
- 1123
- Hard prompts
- 1145
- Long queries
- 1056
- Multi-turn chat
- 1111
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.
-
147 mistral-7b-instruct-v0.2 mistral · Apache-2.0 7B 7GB est +
- Chat Elo (human votes)
- 1149(1142-1155)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 19,402
- Size
- 7B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1185
- Math
- 1127
- Creative writing
- 1104
- Instruction following
- 1122
- Hard prompts
- 1155
- Long queries
- 1133
- Multi-turn chat
- 1112
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
148 wizardlm-13b Microsoft · Llama 2 Community 13B 9GB est +
- Chat Elo (human votes)
- 1149(1139-1158)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 7,044
- Size
- 13B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1151
- Math
- 1064
- Creative writing
- 1140
- Instruction following
- 1123
- Hard prompts
- 1114
- Long queries
- 1149
- Multi-turn chat
- 1109
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~11 GB at Q4, ~9 GB at an aggressive dynamic quant.
-
149 qwen1.5-7b-chat Alibaba · Qianwen LICENSE 7B 3GB +
- Chat Elo (human votes)
- 1143(1133-1153)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 4,737
- Size
- 7B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1209
- Math
- 1121
- Creative writing
- 1081
- Instruction following
- 1123
- Hard prompts
- 1150
- Long queries
- 1163
- Multi-turn chat
- 1112
RAM per quant (real file sizes, via bartowski)
Q2_K 3GB Q3_K_M 4GB Q4_K_M 5GB Q5_K_M 6GB Q6_K 6GB Q8_0 8GB -
150 phi-3-mini-4k-instruct-june-2024 Microsoft · MIT ? ? +
- Chat Elo (human votes)
- 1142(1136-1149)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 12,297
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1196
- Math
- 1193
- Creative writing
- 1096
- Instruction following
- 1123
- Hard prompts
- 1174
- Long queries
- 1105
- Multi-turn chat
- 1098
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
151 llama-2-13b-chat Meta · Llama 2 Community 13B 9GB est +
- Chat Elo (human votes)
- 1141(1134-1148)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 19,174
- Size
- 13B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1162
- Math
- 1110
- Creative writing
- 1091
- Instruction following
- 1109
- Hard prompts
- 1138
- Long queries
- 1139
- Multi-turn chat
- 1099
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~11 GB at Q4, ~9 GB at an aggressive dynamic quant.
-
152 qwen-14b-chat Alibaba · Qianwen LICENSE 14B 10GB est +
- Chat Elo (human votes)
- 1138(1127-1149)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 4,964
- Size
- 14B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 1197
- Math
- 1126
- Creative writing
- 1102
- Instruction following
- 1117
- Hard prompts
- 1141
- Long queries
- 1126
- Multi-turn chat
- 1095
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~12 GB at Q4, ~10 GB at an aggressive dynamic quant.
-
153 gemma-7b-it Google · Gemma license 7B 7GB est +
- Chat Elo (human votes)
- 1137(1127-1146)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 8,925
- Size
- 7B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1167
- Math
- 1118
- Creative writing
- 1101
- Instruction following
- 1103
- Hard prompts
- 1153
- Long queries
- 1111
- Multi-turn chat
- 1041
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
154 codellama-34b-instruct Meta · Llama 2 Community 34B 18GB est +
- Chat Elo (human votes)
- 1136(1127-1145)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 7,366
- Size
- 34B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 24GB GPU
Benchmarks by category (LMArena Elo)
- Coding
- 1159
- Math
- 1109
- Creative writing
- 1086
- Instruction following
- 1103
- Hard prompts
- 1132
- Long queries
- 1097
- Multi-turn chat
- 1073
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~23 GB at Q4, ~18 GB at an aggressive dynamic quant.
-
155 phi-3-mini-128k-instruct Microsoft · MIT ? ? +
- Chat Elo (human votes)
- 1129(1121-1136)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 20,685
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- unknown
Benchmarks by category (LMArena Elo)
- Coding
- 1154
- Math
- 1139
- Creative writing
- 1085
- Instruction following
- 1099
- Hard prompts
- 1129
- Long queries
- 1071
- Multi-turn chat
- 1057
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
156 phi-3-mini-4k-instruct Microsoft · MIT ? 1GB +
- Chat Elo (human votes)
- 1127(1121-1134)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 20,118
- Size
- unconfirmed
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1187
- Math
- 1150
- Creative writing
- 1073
- Instruction following
- 1112
- Hard prompts
- 1152
- Long queries
- 1107
- Multi-turn chat
- 1064
RAM per quant (real file sizes, via bartowski)
Q2_K 1GB Q3_K_M 2GB IQ4_XS 2GB Q4_K_M 2GB Q5_K_M 3GB Q6_K 3GB Q8_0 4GB -
157 codellama-70b-instruct Meta · Llama 2 Community 70B 32GB est +
- Chat Elo (human votes)
- 1118(1100-1137)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 1,143
- Size
- 70B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 32GB
Benchmarks by category (LMArena Elo)
- Instruction following
- 1096
- Hard prompts
- 1159
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.
-
158 gemma-1.1-2b-it Google · Gemma license 2B 1GB +
- Chat Elo (human votes)
- 1116(1108-1123)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 10,854
- Size
- 2B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1171
- Math
- 1109
- Creative writing
- 1093
- Instruction following
- 1099
- Hard prompts
- 1138
- Long queries
- 1116
- Multi-turn chat
- 1050
RAM per quant (real file sizes, via lmstudio-community)
Q2_K 1GB Q3_K_M 1GB IQ4_XS 2GB Q4_K_M 2GB Q5_K_M 2GB Q6_K 2GB Q8_0 3GB -
159 llama-3.2-1b-instruct Meta · Llama 3.2 1B 1GB +
- Chat Elo (human votes)
- 1111(1103-1118)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 8,045
- Size
- 1B
- Context
- 60K
- Multimodal
- text only
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1148
- Math
- 1124
- Creative writing
- 1082
- Instruction following
- 1085
- Hard prompts
- 1113
- Long queries
- 1101
- Multi-turn chat
- 1075
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 1GB UD-IQ1_M 1GB UD-IQ2_XXS 1GB UD-IQ2_M 1GB Q2_K 1GB UD-Q2_K_XL 1GB Q3_K_M 1GB UD-Q3_K_XL 1GB IQ4_XS 1GB Q4_K_M 1GB UD-Q4_K_XL 1GB Q5_K_M 1GB UD-Q5_K_XL 1GB Q6_K 1GB Q8_0 1GB -
160 mistral-7b-instruct mistral · Apache 2.0 7B 7GB est +
- Chat Elo (human votes)
- 1109(1100-1119)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 8,977
- Size
- 7B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1144
- Math
- 1082
- Creative writing
- 1091
- Instruction following
- 1085
- Hard prompts
- 1113
- Long queries
- 1097
- Multi-turn chat
- 1079
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
161 llama-2-7b-chat Meta · Llama 2 Community 7B 7GB est +
- Chat Elo (human votes)
- 1107(1100-1114)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 14,148
- Size
- 7B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1102
- Math
- 1086
- Creative writing
- 1072
- Instruction following
- 1069
- Hard prompts
- 1096
- Long queries
- 1075
- Multi-turn chat
- 1076
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
162 gemma-2b-it Google · Gemma license 2B 5GB est +
- Chat Elo (human votes)
- 1093(1081-1104)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 4,780
- Size
- 2B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1136
- Math
- 1071
- Creative writing
- 1078
- Instruction following
- 1069
- Hard prompts
- 1111
- Long queries
- 1083
- Multi-turn chat
- 1034
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~5 GB at Q4, ~5 GB at an aggressive dynamic quant.
-
163 qwen1.5-4b-chat Alibaba · Qianwen LICENSE 4B 6GB est +
- Chat Elo (human votes)
- 1090(1081-1099)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 7,597
- Size
- 4B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1131
- Math
- 1086
- Creative writing
- 1049
- Instruction following
- 1068
- Hard prompts
- 1094
- Long queries
- 1077
- Multi-turn chat
- 1054
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~6 GB at Q4, ~6 GB at an aggressive dynamic quant.
-
164 olmo-7b-instruct Allen AI · Apache-2.0 7B 7GB est +
- Chat Elo (human votes)
- 1073(1062-1084)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 6,328
- Size
- 7B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 8GB
Benchmarks by category (LMArena Elo)
- Coding
- 1106
- Math
- 1054
- Creative writing
- 1002
- Instruction following
- 1029
- Hard prompts
- 1067
- Multi-turn chat
- 1040
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.
-
165 llama-13b Meta · Non-commercial 13B 9GB est +
- Chat Elo (human votes)
- 973(958-989)
- AA Intelligence (benchmark)
- n/a
- WebDev Elo (coding)
- n/a
- Arena votes
- 2,391
- Size
- 13B
- Context
- n/a
- Multimodal
- n/a
- Smallest rig that fits
- 16GB
Benchmarks by category (LMArena Elo)
- Coding
- 882
- Math
- 921
- Creative writing
- 934
- Instruction following
- 916
- Hard prompts
- 917
- Multi-turn chat
- 891
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~11 GB at Q4, ~9 GB at an aggressive dynamic quant.
-
kimi-k3 Moonshot AI · open weights ? ? +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 57.1 · code 76.2
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- unconfirmed
- Context
- 1M
- Multimodal
- yes (images)
- Smallest rig that fits
- unknown
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
glm-5.2 Zhipu · open weights 744B 217GB +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 51.1 · code 68.8
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- 744B (40B active)
- Context
- 1M
- Multimodal
- text only
- Smallest rig that fits
- 256GB
Zhipu's newest, and the strongest open model out right now on benchmarks. A 744B mixture-of-experts with only 40B active, so it punches like a giant but runs lighter than its size. Needs a big-RAM machine, but an Unsloth 2-bit dynamic quant squeezes it into about 245GB, so a 256GB Mac can run it. The one to watch. Best for: Top-end local reasoning, coding, agents.
RAM per quant (real file sizes, via unsloth)
UD-IQ1_S 217GB UD-IQ1_M 228GB UD-IQ2_XXS 238GB UD-IQ2_M 239GB UD-Q2_K_XL 254GB UD-Q3_K_XL 343GB UD-Q4_K_XL 467GB UD-Q5_K_XL 562GB Q8_0 801GB -
kimi-k2.7-code Moonshot AI · open weights ? 304GB +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 41.9 · code 60.8
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- unconfirmed
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 512GB Mac Studio
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 304GB UD-IQ2_XXS 318GB UD-IQ2_M 318GB UD-Q2_K_XL 339GB UD-Q3_K_XL 464GB UD-Q4_K_XL 584GB -
hy3-preview Tencent · open weights ? ? +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 41.2 · code 58.8
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- unconfirmed
- Context
- 262K
- Multimodal
- text only
- Smallest rig that fits
- unknown
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
nex-n2-pro nex-agi · open weights ? ? +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 41.0 · code 59.1
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- unconfirmed
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- unknown
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
-
nemotron-3-ultra-550b-a55b NVIDIA · open weights 550B 224GB est +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 37.8 · code 49.3
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- 550B (55B active)
- Context
- 512K
- Multimodal
- text only
- Smallest rig that fits
- 256GB
RAM per quant
No trusted GGUF build found yet. Estimated from size: ~307 GB at Q4, ~224 GB at an aggressive dynamic quant.
-
qwen3.6-27b Alibaba · open weights 27B 10GB +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 37.1 · code 53.7
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- 27B
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 16GB
RAM per quant (real file sizes, via unsloth)
UD-IQ2_XXS 10GB UD-IQ2_M 11GB UD-Q2_K_XL 12GB Q3_K_M 14GB UD-Q3_K_XL 15GB IQ4_XS 16GB Q4_K_M 17GB UD-Q4_K_XL 18GB Q5_K_M 20GB UD-Q5_K_XL 20GB Q6_K 23GB Q8_0 29GB -
qwen3.6-35b-a3b Alibaba · open weights 35B 10GB +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 31.6 · code 41.9
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- 35B (3B active)
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 16GB
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 10GB UD-IQ2_XXS 11GB UD-IQ2_M 12GB UD-Q2_K_XL 12GB UD-Q3_K_XL 17GB UD-Q4_K_XL 22GB UD-Q5_K_XL 27GB Q8_0 37GB -
step-3.7-flash stepfun · open weights ? 57GB +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 30.3 · code 39.6
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- unconfirmed
- Context
- 262K
- Multimodal
- yes (images)
- Smallest rig that fits
- 64GB Mac
RAM per quant (real file sizes, via unsloth)
UD-IQ1_M 57GB UD-IQ2_XXS 62GB UD-IQ2_M 62GB UD-Q2_K_XL 66GB UD-Q3_K_XL 89GB UD-Q4_K_XL 122GB UD-Q5_K_XL 146GB Q8_0 209GB -
trinity-large-thinking arcee-ai · open weights ? ? +
- Chat Elo (human votes)
- not yet rated
- AA Intelligence (benchmark)
- 18.2 · code 25.8
- WebDev Elo (coding)
- n/a
- Arena votes
- n/a
- Size
- unconfirmed
- Context
- 262K
- Multimodal
- text only
- Smallest rig that fits
- unknown
RAM per quant
No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.
Showing the current top of the field. Older, lower-ranked releases are hidden by default.
No model matches that search and hardware combination. Clear the search, or try a bigger rig.
How we rank
Two rankings from two sources. Human votes: LMArena style-controlled Elo, which adjusts for response length and formatting, from blind human-preference votes. Benchmarks: Artificial Analysis Intelligence (republished via OpenRouter), which scores models within days of release, so it catches models the arena has not voted on yet. RAM is taken from real GGUF file sizes where available, otherwise estimated from parameter count.
- Two rankings, two sources. Human votes is LMArena style-controlled Elo, from millions of blind human-preference votes: the gold standard, but a new model needs weeks of votes before it gets an Elo. Benchmarks is Artificial Analysis Intelligence (republished by OpenRouter), a test-based score that lands within days of a launch. Toggle between them. The newest models, the ones the arena has not voted on yet, are pulled in from OpenRouter and shown with their benchmark score and a "new" tag.
- Other arena scores: WebDev Elo is LMArena's web-development arena (the closest open proxy for coding). Category numbers (Coding, Math, Creative writing, and so on) are LMArena's own category-filtered Elo, from the same votes, so you can pick a model for a specific job.
- Category benchmarks: the per-category numbers (Coding, Math, Creative writing, Instruction following, Hard prompts, Long queries, Multi-turn) are LMArena's own category-filtered Elo for that model, from the same human votes. They tell you what a model is relatively good at; a model can rank mid-table overall but punch above its weight in coding or math. The main ranking always uses the overall number.
- RAM per quant is real: where a trusted publisher (Unsloth, llama.cpp, LM Studio, bartowski) ships a GGUF build, we read the actual file size of every quant from Hugging Face. File size is, in practice, the RAM you need, plus a few GB of headroom for context.
- Strict matching: a model only gets quant data when its name matches the GGUF repo exactly. "No trusted GGUF build found yet" means exactly that, not a guess at a lookalike. New builds get picked up automatically by the daily refresh.
- Estimates are labeled: where no GGUF exists, RAM is estimated from parameter count (~0.55 bytes/param at Q4, ~0.4 dynamic) and marked "est".
- Context window and modality come from OpenRouter's live model data, attached only on an exact match.
Sources: LMArena (human votes), Artificial Analysis (benchmarks, via OpenRouter), Hugging Face (quant file sizes). Parameter counts and licenses are from the model providers' own cards.
Open-source LLM FAQ
What is the best open-source LLM right now? +
It depends which signal you trust. By human votes (LMArena Elo) the leader as of July 27, 2026 is glm-5.1 from Zhipu.
Why is a brand-new model not in the arena ranking? +
LMArena Elo is computed from thousands of blind human votes, which take days to weeks to accumulate after a model launches. So a model released this week has no Elo yet. We catch those releases from OpenRouter and show them in the "Just released" section and the Benchmarks ranking with their Artificial Analysis score, until the arena catches up.
How much RAM do I need to run an open LLM at home? +
The quant file size is, in practice, the RAM you need, plus a few GB of headroom for context. Where a trusted GGUF build exists we list the real file size of every quant. As a rule of thumb, Q4 (the everyday sweet spot) is roughly half the parameter count in GB.
Where can I download open LLM weights? +
Expand any model on this page. Where a canonical GGUF build exists we link it directly, plus Hugging Face GGUF search, MLX builds for Apple Silicon, and Ollama.
What do the category benchmark numbers mean? +
Each expanded model shows LMArena Elo filtered by task category: coding, math, creative writing, instruction following, hard prompts, long queries, and multi-turn chat. They come from the same blind human votes as the main ranking, just sliced by what the conversation was about. Use them to pick a model for a specific job, e.g. the best coder that fits your RAM rather than the best all-rounder.
Which quant should I download? +
Q4_K_M is the everyday default: most of the quality at about a quarter of the full size. If the model is just over your RAM budget, try an Unsloth dynamic quant (UD-), which holds quality better at smaller sizes. Q8_0 is near-lossless if you have RAM to spare.
How is this list ranked and updated? +
Models are ranked by LMArena style-controlled Elo from millions of blind human votes, filtered to open-weight models only. Rankings, quant file sizes, and metadata refresh automatically every day. Last refresh: July 27, 2026.
Want this kind of clarity for your business?
BlueFort AI is the content side of BlueFort IT. When you're ready to actually deploy AI, openly or privately, that's their day job.