Skip to content
Open-source LLM rankings

Every open-source LLM, ranked daily.

Every open-weight model worth knowing, ranked two ways: by human votes (LMArena Elo) and by benchmarks (Artificial Analysis), so the newest models show up the day they launch instead of weeks later. With the exact RAM each quant needs (real file sizes) and where to download it. As of July 27, 2026, the human-vote leader is glm-5.1.

Updated July 27, 2026 · Sources: LMArena style-controlled Elo + Artificial Analysis · quant sizes from real GGUF files · 165 rated + 10 new

165 models
  1. 1 glm-5.1 Zhipu · MIT 206GB
    Chat Elo (human votes)
    1470(1465-1474)
    AA Intelligence (benchmark)
    40.2 · code 55.8
    WebDev Elo (coding)
    1520
    Arena votes
    30,726
    Size
    355B (32B active)
    Context
    205K
    Multimodal
    text only
    Smallest rig that fits
    256GB

    Zhipu's flagship. Excellent all-rounder, strong coding, clean MIT license. Best for: Coding and general agent work.

    Benchmarks by category (LMArena Elo)

    Coding
    1520
    Math
    1480
    Creative writing
    1454
    Instruction following
    1464
    Hard prompts
    1492
    Long queries
    1484
    Multi-turn chat
    1483

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 206GB UD-IQ2_XXS 221GB UD-IQ2_M 236GB UD-Q2_K_XL 252GB UD-Q3_K_XL 340GB UD-Q4_K_XL 466GB UD-Q5_K_XL 560GB Q8_0 801GB
  2. 2 glm-5.2 (max) Zhipu · MIT ?
    Chat Elo (human votes)
    1469(1463-1475)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1588
    Arena votes
    18,017
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1509
    Math
    1474
    Creative writing
    1446
    Instruction following
    1463
    Hard prompts
    1490
    Long queries
    1479
    Multi-turn chat
    1471

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  3. 3 mimo-v2.5-pro Xiaomi · MIT 304GB
    Chat Elo (human votes)
    1467(1462-1471)
    AA Intelligence (benchmark)
    42.2 · code 60.2
    WebDev Elo (coding)
    1476
    Arena votes
    42,195
    Size
    unconfirmed
    Context
    1M
    Multimodal
    text only
    Smallest rig that fits
    512GB Mac Studio

    Xiaomi's flagship MiMo. Strong scores; confirm size on the model card. Best for: General use (verify specs).

    Benchmarks by category (LMArena Elo)

    Coding
    1520
    Math
    1475
    Creative writing
    1433
    Instruction following
    1470
    Hard prompts
    1495
    Long queries
    1488
    Multi-turn chat
    1477

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 304GB UD-IQ2_XXS 317GB UD-IQ2_M 317GB UD-Q2_K_XL 338GB UD-Q3_K_XL 460GB UD-Q4_K_XL 631GB UD-Q5_K_XL 759GB Q8_0 1088GB
  4. 4 kimi-k2.6 Moonshot AI · Modified MIT 340GB
    Chat Elo (human votes)
    1461(1456-1465)
    AA Intelligence (benchmark)
    44.2 · code 61.8
    WebDev Elo (coding)
    1510
    Arena votes
    37,686
    Size
    1000B (32B active)
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    512GB Mac Studio

    Frontier-class open reasoning. Huge MoE, so it needs serious memory, but on a big-RAM Mac it flies. Best for: Top-end local reasoning and agents.

    Benchmarks by category (LMArena Elo)

    Coding
    1515
    Math
    1480
    Creative writing
    1430
    Instruction following
    1455
    Hard prompts
    1485
    Long queries
    1476
    Multi-turn chat
    1460

    RAM per quant (real file sizes, via unsloth)

    UD-Q2_K_XL 340GB UD-Q4_K_XL 584GB
  5. 5 glm-5 Zhipu · MIT 204GB
    Chat Elo (human votes)
    1457(1452-1461)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1435
    Arena votes
    27,785
    Size
    355B (32B active)
    Context
    205K
    Multimodal
    text only
    Smallest rig that fits
    256GB

    The prior GLM-5. Still elite, often cheaper to run than 5.1. Best for: General use, coding.

    Benchmarks by category (LMArena Elo)

    Coding
    1497
    Math
    1443
    Creative writing
    1445
    Instruction following
    1447
    Hard prompts
    1478
    Long queries
    1470
    Multi-turn chat
    1472

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 204GB UD-IQ1_M 224GB UD-IQ2_XXS 241GB UD-IQ2_M 255GB Q2_K 276GB UD-Q2_K_XL 281GB Q3_K_M 360GB UD-Q3_K_XL 332GB IQ4_XS 403GB Q4_K_M 456GB UD-Q4_K_XL 431GB Q5_K_M 535GB UD-Q5_K_XL 536GB Q6_K 619GB Q8_0 801GB
  6. 6 deepseek-v4-pro DeepSeek · MIT 272GB est
    Chat Elo (human votes)
    1457(1452-1461)
    AA Intelligence (benchmark)
    44.3 · code 59.4
    WebDev Elo (coding)
    1447
    Arena votes
    45,278
    Size
    671B (37B active)
    Context
    1M
    Multimodal
    text only
    Smallest rig that fits
    512GB Mac Studio

    DeepSeek's big MoE. Exceptional at code and math, permissive MIT. Best for: Coding, math, large-context work.

    Benchmarks by category (LMArena Elo)

    Coding
    1501
    Math
    1445
    Creative writing
    1442
    Instruction following
    1453
    Hard prompts
    1480
    Long queries
    1473
    Multi-turn chat
    1472

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~373 GB at Q4, ~272 GB at an aggressive dynamic quant.

  7. 7 deepseek-v4-pro-thinking DeepSeek · MIT 272GB est
    Chat Elo (human votes)
    1455(1451-1460)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1464
    Arena votes
    43,079
    Size
    671B (37B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    512GB Mac Studio

    Reasoning mode of V4 Pro. Best DeepSeek for hard, multi-step problems. Best for: Hard reasoning and coding.

    Benchmarks by category (LMArena Elo)

    Coding
    1490
    Math
    1467
    Creative writing
    1443
    Instruction following
    1447
    Hard prompts
    1475
    Long queries
    1466
    Multi-turn chat
    1456

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~373 GB at Q4, ~272 GB at an aggressive dynamic quant.

  8. 8 gemma-4-31b Google · Apache 2.0 9GB
    Chat Elo (human votes)
    1451(1443-1458)
    AA Intelligence (benchmark)
    29.4 · code 43.4
    WebDev Elo (coding)
    1364
    Arena votes
    5,880
    Size
    31B
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    16GB

    Google's dense 31B. The sweet spot for a single 24GB GPU. Punches above its size. Best for: Single-GPU and laptops.

    Benchmarks by category (LMArena Elo)

    Coding
    1498
    Math
    1470
    Creative writing
    1421
    Instruction following
    1452
    Hard prompts
    1473
    Long queries
    1467
    Multi-turn chat
    1464

    RAM per quant (real file sizes, via unsloth)

    UD-IQ2_XXS 9GB UD-IQ2_M 11GB UD-Q2_K_XL 12GB Q3_K_M 15GB UD-Q3_K_XL 15GB IQ4_XS 16GB Q4_K_M 18GB UD-Q4_K_XL 19GB Q5_K_M 22GB UD-Q5_K_XL 22GB Q6_K 25GB Q8_0 33GB
  9. 9 kimi-k2.5-thinking Moonshot AI · Modified MIT 276GB
    Chat Elo (human votes)
    1450(1446-1453)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1437
    Arena votes
    62,688
    Size
    1000B (32B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    512GB Mac Studio

    The deliberate, chain-of-thought Kimi. Slower, stronger on hard problems. Best for: Complex reasoning, planning.

    Benchmarks by category (LMArena Elo)

    Coding
    1502
    Math
    1472
    Creative writing
    1424
    Instruction following
    1440
    Hard prompts
    1471
    Long queries
    1458
    Multi-turn chat
    1452

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 276GB UD-IQ1_M 301GB UD-IQ2_XXS 327GB UD-IQ2_M 345GB Q2_K 374GB UD-Q2_K_XL 375GB Q3_K_M 490GB UD-Q3_K_XL 490GB IQ4_XS 547GB Q4_K_M 621GB UD-Q4_K_XL 622GB Q5_K_M 729GB UD-Q5_K_XL 731GB Q6_K 843GB Q8_0 1091GB
  10. 10 inkling thinky · Apache 2.0 270GB
    Chat Elo (human votes)
    1445(1437-1453)
    AA Intelligence (benchmark)
    40.7 · code 52.1
    WebDev Elo (coding)
    1418
    Arena votes
    5,386
    Size
    unconfirmed
    Context
    1M
    Multimodal
    yes (images)
    Smallest rig that fits
    512GB Mac Studio

    Benchmarks by category (LMArena Elo)

    Coding
    1498
    Math
    1476
    Creative writing
    1395
    Instruction following
    1428
    Hard prompts
    1470
    Long queries
    1448
    Multi-turn chat
    1454

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 270GB UD-IQ1_M 285GB UD-Q2_K_XL 317GB UD-Q3_K_XL 433GB UD-Q4_K_XL 587GB Q8_0 857GB
  11. 11 minimax-m3 minimax · MiniMax Community License 128GB
    Chat Elo (human votes)
    1444(1439-1449)
    AA Intelligence (benchmark)
    44.4 · code 58.6
    WebDev Elo (coding)
    1493
    Arena votes
    28,093
    Size
    unconfirmed
    Context
    1M
    Multimodal
    yes (images)
    Smallest rig that fits
    128GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1498
    Math
    1443
    Creative writing
    1407
    Instruction following
    1437
    Hard prompts
    1464
    Long queries
    1453
    Multi-turn chat
    1452

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 128GB UD-IQ2_XXS 134GB UD-IQ2_M 134GB UD-Q2_K_XL 143GB UD-Q3_K_XL 195GB UD-Q4_K_XL 265GB UD-Q5_K_XL 318GB Q8_0 453GB
  12. 12 qwen3.5-397b-a17b Alibaba · Apache 2.0 107GB
    Chat Elo (human votes)
    1442(1438-1446)
    AA Intelligence (benchmark)
    33.7 · code 48.2
    WebDev Elo (coding)
    1400
    Arena votes
    58,157
    Size
    397B (17B active)
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    128GB Mac

    Alibaba's big sparse MoE: 397B total but only 17B active, so it runs faster than its size suggests. Apache 2.0. Best for: Best license, multilingual, agents.

    Benchmarks by category (LMArena Elo)

    Coding
    1492
    Math
    1446
    Creative writing
    1409
    Instruction following
    1433
    Hard prompts
    1463
    Long queries
    1455
    Multi-turn chat
    1451

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 107GB UD-IQ2_XXS 115GB UD-IQ2_M 123GB Q3_K_M 177GB UD-Q3_K_XL 179GB Q4_K_M 244GB UD-Q4_K_XL 245GB Q5_K_M 294GB UD-Q5_K_XL 295GB Q6_K 327GB Q8_0 422GB
  13. 13 glm-4.7 Zhipu · MIT 97GB
    Chat Elo (human votes)
    1442(1436-1448)
    AA Intelligence (benchmark)
    33.7 · code 45.3
    WebDev Elo (coding)
    1433
    Arena votes
    12,092
    Size
    355B (32B active)
    Context
    205K
    Multimodal
    text only
    Smallest rig that fits
    128GB Mac

    Last-gen GLM. A proven, stable workhorse if you want maturity over bleeding edge. Best for: Reliable daily driver.

    Benchmarks by category (LMArena Elo)

    Coding
    1485
    Math
    1428
    Creative writing
    1405
    Instruction following
    1428
    Hard prompts
    1463
    Long queries
    1452
    Multi-turn chat
    1459

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 97GB UD-IQ1_M 108GB UD-IQ2_XXS 116GB UD-IQ2_M 122GB Q2_K 131GB UD-Q2_K_XL 135GB Q3_K_M 171GB UD-Q3_K_XL 159GB IQ4_XS 192GB Q4_K_M 216GB UD-Q4_K_XL 205GB Q5_K_M 254GB UD-Q5_K_XL 254GB Q6_K 294GB Q8_0 381GB
  14. 14 deepseek-v4-flash-thinking DeepSeek · MIT 84GB est
    Chat Elo (human votes)
    1439(1434-1443)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    44,822
    Size
    200B (20B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    96GB

    The lighter, faster DeepSeek. Easier to run, still a strong reasoner. Best for: Reasoning on mid-range hardware.

    Benchmarks by category (LMArena Elo)

    Coding
    1481
    Math
    1440
    Creative writing
    1406
    Instruction following
    1436
    Hard prompts
    1460
    Long queries
    1451
    Multi-turn chat
    1448

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~114 GB at Q4, ~84 GB at an aggressive dynamic quant.

  15. 15 gemma-4-26b-a4b Google · Apache 2.0 10GB
    Chat Elo (human votes)
    1438(1430-1446)
    AA Intelligence (benchmark)
    25.7 · code 39.3
    WebDev Elo (coding)
    1366
    Arena votes
    5,798
    Size
    26B (4B active)
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1480
    Math
    1466
    Creative writing
    1402
    Instruction following
    1439
    Hard prompts
    1461
    Long queries
    1448
    Multi-turn chat
    1446

    RAM per quant (real file sizes, via unsloth)

    UD-IQ2_XXS 10GB UD-IQ2_M 10GB UD-Q2_K_XL 11GB UD-Q3_K_XL 13GB UD-Q4_K_XL 17GB UD-Q5_K_XL 21GB Q8_0 27GB
  16. 16 deepseek-v4-flash DeepSeek · MIT 83GB
    Chat Elo (human votes)
    1436(1431-1440)
    AA Intelligence (benchmark)
    40.3 · code 56.2
    WebDev Elo (coding)
    n/a
    Arena votes
    45,050
    Size
    unconfirmed
    Context
    1M
    Multimodal
    text only
    Smallest rig that fits
    96GB

    Benchmarks by category (LMArena Elo)

    Coding
    1482
    Math
    1426
    Creative writing
    1408
    Instruction following
    1428
    Hard prompts
    1459
    Long queries
    1449
    Multi-turn chat
    1453

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 83GB UD-IQ1_M 87GB UD-IQ2_XXS 91GB UD-IQ2_M 91GB UD-Q2_K_XL 97GB UD-Q3_K_XL 129GB UD-Q4_K_XL 155GB
  17. 17 mimo-v2.5 Xiaomi · MIT 93GB
    Chat Elo (human votes)
    1433(1429-1437)
    AA Intelligence (benchmark)
    37.2 · code 56.8
    WebDev Elo (coding)
    1435
    Arena votes
    43,091
    Size
    unconfirmed
    Context
    1M
    Multimodal
    yes (images)
    Smallest rig that fits
    96GB

    The standard MiMo 2.5. Lighter than Pro; check the card for exact size. Best for: General use (verify specs).

    Benchmarks by category (LMArena Elo)

    Coding
    1491
    Math
    1441
    Creative writing
    1392
    Instruction following
    1431
    Hard prompts
    1461
    Long queries
    1451
    Multi-turn chat
    1449

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 93GB UD-IQ2_XXS 96GB UD-IQ2_M 97GB UD-Q2_K_XL 103GB UD-Q3_K_XL 140GB UD-Q4_K_XL 192GB UD-Q5_K_XL 231GB Q8_0 329GB
  18. 18 kimi-k2.5-instant Moonshot AI · Modified MIT 276GB
    Chat Elo (human votes)
    1431(1425-1438)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1405
    Arena votes
    8,178
    Size
    1000B (32B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    512GB Mac Studio

    Kimi tuned for speed over deep thinking. Snappy for chat and tools. Best for: Fast local assistant.

    Benchmarks by category (LMArena Elo)

    Coding
    1504
    Math
    1441
    Creative writing
    1390
    Instruction following
    1435
    Hard prompts
    1461
    Long queries
    1445
    Multi-turn chat
    1439

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 276GB UD-IQ1_M 301GB UD-IQ2_XXS 327GB UD-IQ2_M 345GB Q2_K 374GB UD-Q2_K_XL 375GB Q3_K_M 490GB UD-Q3_K_XL 490GB IQ4_XS 547GB Q4_K_M 621GB UD-Q4_K_XL 622GB Q5_K_M 729GB UD-Q5_K_XL 731GB Q6_K 843GB Q8_0 1091GB
  19. 19 kimi-k2-thinking-turbo Moonshot AI · Modified MIT 280GB
    Chat Elo (human votes)
    1430(1427-1433)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1321
    Arena votes
    61,950
    Size
    1000B (32B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    512GB Mac Studio

    Faster thinking variant of the previous Kimi gen. Still very capable. Best for: Reasoning on a tighter time budget.

    Benchmarks by category (LMArena Elo)

    Coding
    1486
    Math
    1438
    Creative writing
    1392
    Instruction following
    1417
    Hard prompts
    1453
    Long queries
    1433
    Multi-turn chat
    1433

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 280GB UD-IQ1_M 304GB UD-IQ2_XXS 329GB UD-IQ2_M 347GB Q2_K 373GB UD-Q2_K_XL 382GB Q3_K_M 489GB UD-Q3_K_XL 452GB IQ4_XS 547GB Q4_K_M 621GB UD-Q4_K_XL 587GB Q5_K_M 728GB UD-Q5_K_XL 731GB Q6_K 843GB Q8_0 1091GB
  20. 20 mistral-medium-3.5 mistral · Modified MIT ?
    Chat Elo (human votes)
    1427(1421-1434)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1267
    Arena votes
    11,019
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1479
    Math
    1433
    Creative writing
    1395
    Instruction following
    1421
    Hard prompts
    1446
    Long queries
    1431
    Multi-turn chat
    1432

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  21. 21 nvidia-nemotron-3-ultra-550b-a55b-nvfp4 NVIDIA · OpenMDW-1.1 224GB est
    Chat Elo (human votes)
    1425(1418-1432)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    10,562
    Size
    550B (55B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    256GB

    Benchmarks by category (LMArena Elo)

    Coding
    1473
    Math
    1446
    Creative writing
    1376
    Instruction following
    1405
    Hard prompts
    1442
    Long queries
    1433
    Multi-turn chat
    1398

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~307 GB at Q4, ~224 GB at an aggressive dynamic quant.

  22. 22 glm-4.6 Zhipu · MIT 97GB
    Chat Elo (human votes)
    1425(1421-1429)
    AA Intelligence (benchmark)
    28.7 · code 45.8
    WebDev Elo (coding)
    1339
    Arena votes
    35,608
    Size
    unconfirmed
    Context
    205K
    Multimodal
    text only
    Smallest rig that fits
    128GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1459
    Math
    1420
    Creative writing
    1402
    Instruction following
    1415
    Hard prompts
    1442
    Long queries
    1433
    Multi-turn chat
    1421

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 97GB UD-IQ1_M 107GB UD-IQ2_XXS 115GB UD-IQ2_M 122GB Q2_K 131GB UD-Q2_K_XL 135GB Q3_K_M 171GB UD-Q3_K_XL 158GB IQ4_XS 191GB Q4_K_M 216GB UD-Q4_K_XL 204GB Q5_K_M 253GB UD-Q5_K_XL 252GB Q6_K 293GB Q8_0 379GB
  23. 23 deepseek-v3.2 DeepSeek · MIT 184GB
    Chat Elo (human votes)
    1425(1421-1428)
    AA Intelligence (benchmark)
    32.0 · code 44.2
    WebDev Elo (coding)
    1323
    Arena votes
    47,206
    Size
    unconfirmed
    Context
    164K
    Multimodal
    text only
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1470
    Math
    1429
    Creative writing
    1401
    Instruction following
    1421
    Hard prompts
    1447
    Long queries
    1442
    Multi-turn chat
    1428

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 184GB UD-IQ1_M 199GB UD-IQ2_XXS 217GB UD-IQ2_M 228GB Q2_K 245GB UD-Q2_K_XL 247GB Q3_K_M 320GB UD-Q3_K_XL 321GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 408GB Q5_K_M 476GB UD-Q5_K_XL 482GB Q6_K 551GB Q8_0 713GB
  24. 24 deepseek-v3.2-exp-thinking DeepSeek · MIT ?
    Chat Elo (human votes)
    1425(1418-1431)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    9,069
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1474
    Math
    1428
    Creative writing
    1394
    Instruction following
    1417
    Hard prompts
    1446
    Long queries
    1431
    Multi-turn chat
    1423

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  25. 25 qwen3-235b-a22b-instruct-2507 Alibaba · Apache 2.0 86GB
    Chat Elo (human votes)
    1423(1421-1426)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    97,085
    Size
    235B (22B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    96GB

    Benchmarks by category (LMArena Elo)

    Coding
    1473
    Math
    1419
    Creative writing
    1380
    Instruction following
    1416
    Hard prompts
    1448
    Long queries
    1435
    Multi-turn chat
    1438

    RAM per quant (real file sizes, via unsloth)

    Q2_K 86GB UD-Q2_K_XL 89GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 125GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 169GB Q6_K 193GB Q8_0 250GB
  26. 26 deepseek-v3.2-thinking DeepSeek · MIT 184GB
    Chat Elo (human votes)
    1423(1419-1427)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1361
    Arena votes
    41,025
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1475
    Math
    1425
    Creative writing
    1391
    Instruction following
    1419
    Hard prompts
    1446
    Long queries
    1442
    Multi-turn chat
    1427

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 184GB UD-IQ1_M 199GB UD-IQ2_XXS 217GB UD-IQ2_M 228GB Q2_K 245GB UD-Q2_K_XL 247GB Q3_K_M 320GB UD-Q3_K_XL 321GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 408GB Q5_K_M 476GB UD-Q5_K_XL 482GB Q6_K 551GB Q8_0 713GB
  27. 27 deepseek-v3.2-exp DeepSeek · MIT ?
    Chat Elo (human votes)
    1423(1416-1429)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1272
    Arena votes
    11,913
    Size
    unconfirmed
    Context
    164K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1465
    Math
    1417
    Creative writing
    1410
    Instruction following
    1416
    Hard prompts
    1447
    Long queries
    1442
    Multi-turn chat
    1430

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  28. 28 deepseek-r1-0528 DeepSeek · MIT 185GB
    Chat Elo (human votes)
    1422(1416-1428)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    18,451
    Size
    unconfirmed
    Context
    164K
    Multimodal
    text only
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1464
    Math
    1396
    Creative writing
    1394
    Instruction following
    1392
    Hard prompts
    1433
    Long queries
    1407
    Multi-turn chat
    1406

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 185GB UD-IQ1_M 200GB UD-IQ2_XXS 217GB UD-IQ2_M 229GB Q2_K 245GB UD-Q2_K_XL 251GB Q3_K_M 320GB UD-Q3_K_XL 296GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 384GB Q5_K_M 476GB UD-Q5_K_XL 481GB Q6_K 551GB Q8_0 713GB
  29. 29 kimi-k2-0905-preview Moonshot AI · Modified MIT ?
    Chat Elo (human votes)
    1418(1411-1424)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    11,770
    Size
    unconfirmed
    Context
    262K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1468
    Math
    1416
    Creative writing
    1382
    Instruction following
    1390
    Hard prompts
    1436
    Long queries
    1402
    Multi-turn chat
    1404

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  30. 30 deepseek-v3.1 DeepSeek · MIT 192GB
    Chat Elo (human votes)
    1417(1411-1423)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    14,943
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1448
    Math
    1415
    Creative writing
    1389
    Instruction following
    1403
    Hard prompts
    1433
    Long queries
    1421
    Multi-turn chat
    1406

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 192GB UD-IQ1_M 207GB UD-IQ2_XXS 226GB UD-IQ2_M 236GB Q2_K 246GB UD-Q2_K_XL 256GB Q3_K_M 320GB UD-Q3_K_XL 300GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 387GB Q5_K_M 476GB UD-Q5_K_XL 485GB Q6_K 551GB Q8_0 713GB
  31. 31 kimi-k2-0711-preview Moonshot AI · Modified MIT ?
    Chat Elo (human votes)
    1417(1413-1422)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    27,608
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1460
    Math
    1388
    Creative writing
    1371
    Instruction following
    1379
    Hard prompts
    1431
    Long queries
    1393
    Multi-turn chat
    1420

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  32. 32 qwen3.5-122b-a10b Alibaba · Apache 2.0 34GB
    Chat Elo (human votes)
    1417(1413-1422)
    AA Intelligence (benchmark)
    32.3 · code 45.7
    WebDev Elo (coding)
    1359
    Arena votes
    28,504
    Size
    122B (10B active)
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    48GB

    Benchmarks by category (LMArena Elo)

    Coding
    1459
    Math
    1424
    Creative writing
    1367
    Instruction following
    1407
    Hard prompts
    1433
    Long queries
    1419
    Multi-turn chat
    1418

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 34GB UD-IQ2_XXS 37GB UD-IQ2_M 39GB UD-Q2_K_XL 42GB Q3_K_M 56GB UD-Q3_K_XL 57GB Q4_K_M 77GB UD-Q4_K_XL 77GB Q5_K_M 92GB UD-Q5_K_XL 92GB Q6_K 101GB Q8_0 130GB
  33. 33 deepseek-v3.1-terminus-thinking DeepSeek · MIT 187GB
    Chat Elo (human votes)
    1417(1407-1427)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    3,457
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1463
    Math
    1406
    Creative writing
    1385
    Instruction following
    1420
    Hard prompts
    1445
    Long queries
    1443
    Multi-turn chat
    1419

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 187GB UD-IQ1_M 201GB UD-IQ2_XXS 217GB UD-IQ2_M 229GB Q2_K 246GB UD-Q2_K_XL 251GB Q3_K_M 320GB UD-Q3_K_XL 296GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 384GB Q5_K_M 476GB UD-Q5_K_XL 482GB Q6_K 551GB Q8_0 713GB
  34. 34 deepseek-v3.1-thinking DeepSeek · MIT 192GB
    Chat Elo (human votes)
    1417(1410-1424)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    11,726
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1457
    Math
    1414
    Creative writing
    1404
    Instruction following
    1419
    Hard prompts
    1437
    Long queries
    1446
    Multi-turn chat
    1414

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 192GB UD-IQ1_M 207GB UD-IQ2_XXS 226GB UD-IQ2_M 236GB Q2_K 246GB UD-Q2_K_XL 256GB Q3_K_M 320GB UD-Q3_K_XL 300GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 387GB Q5_K_M 476GB UD-Q5_K_XL 485GB Q6_K 551GB Q8_0 713GB
  35. 35 minimax-m2.7 minimax · Modified MIT 61GB
    Chat Elo (human votes)
    1417(1413-1421)
    AA Intelligence (benchmark)
    38.1 · code 52.6
    WebDev Elo (coding)
    1398
    Arena votes
    49,930
    Size
    unconfirmed
    Context
    205K
    Multimodal
    text only
    Smallest rig that fits
    64GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1479
    Math
    1424
    Creative writing
    1365
    Instruction following
    1410
    Hard prompts
    1443
    Long queries
    1435
    Multi-turn chat
    1428

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 61GB UD-IQ2_XXS 65GB UD-IQ2_M 70GB UD-Q2_K_XL 75GB UD-Q3_K_XL 102GB UD-Q4_K_XL 141GB UD-Q5_K_XL 169GB Q8_0 243GB
  36. 36 mistral-large-3 mistral · Apache 2.0 ?
    Chat Elo (human votes)
    1415(1412-1419)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1230
    Arena votes
    50,849
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1467
    Math
    1402
    Creative writing
    1374
    Instruction following
    1403
    Hard prompts
    1433
    Long queries
    1417
    Multi-turn chat
    1421

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  37. 37 deepseek-v3.1-terminus DeepSeek · MIT 187GB
    Chat Elo (human votes)
    1415(1406-1425)
    AA Intelligence (benchmark)
    30.4 · code 43.5
    WebDev Elo (coding)
    n/a
    Arena votes
    3,692
    Size
    unconfirmed
    Context
    164K
    Multimodal
    text only
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1439
    Math
    1395
    Creative writing
    1407
    Instruction following
    1394
    Hard prompts
    1424
    Long queries
    1418
    Multi-turn chat
    1394

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 187GB UD-IQ1_M 201GB UD-IQ2_XXS 217GB UD-IQ2_M 229GB Q2_K 246GB UD-Q2_K_XL 251GB Q3_K_M 320GB UD-Q3_K_XL 296GB IQ4_XS 358GB Q4_K_M 405GB UD-Q4_K_XL 384GB Q5_K_M 476GB UD-Q5_K_XL 482GB Q6_K 551GB Q8_0 713GB
  38. 38 qwen3-vl-235b-a22b-instruct Alibaba · Apache 2.0 63GB
    Chat Elo (human votes)
    1415(1409-1422)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    11,499
    Size
    235B (22B active)
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    64GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1465
    Math
    1410
    Creative writing
    1361
    Instruction following
    1414
    Hard prompts
    1440
    Long queries
    1424
    Multi-turn chat
    1426

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 63GB UD-IQ1_M 70GB Q2_K 86GB UD-Q2_K_XL 88GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 125GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 167GB Q6_K 193GB Q8_0 250GB
  39. 39 hunyuan-hy3-preview Tencent · tencent-hunyuan-community 160GB est
    Chat Elo (human votes)
    1412(1405-1420)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1357
    Arena votes
    6,634
    Size
    389B (52B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Tencent's Hunyuan preview. Capable MoE; preview, so expect changes. Best for: Experimentation.

    Benchmarks by category (LMArena Elo)

    Coding
    1459
    Math
    1429
    Creative writing
    1356
    Instruction following
    1399
    Hard prompts
    1438
    Long queries
    1425
    Multi-turn chat
    1415

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~218 GB at Q4, ~160 GB at an aggressive dynamic quant.

  40. 40 glm-4.5 Zhipu · MIT 97GB
    Chat Elo (human votes)
    1411(1406-1416)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    24,286
    Size
    unconfirmed
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    128GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1454
    Math
    1413
    Creative writing
    1373
    Instruction following
    1405
    Hard prompts
    1433
    Long queries
    1416
    Multi-turn chat
    1406

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 97GB UD-IQ1_M 108GB UD-IQ2_XXS 116GB UD-IQ2_M 122GB Q2_K 131GB UD-Q2_K_XL 135GB Q3_K_M 171GB UD-Q3_K_XL 159GB IQ4_XS 192GB Q4_K_M 216GB UD-Q4_K_XL 204GB Q5_K_M 254GB UD-Q5_K_XL 253GB Q6_K 294GB Q8_0 381GB
  41. 41 qwen3.5-27b Alibaba · Apache 2.0 9GB
    Chat Elo (human votes)
    1409(1404-1413)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1357
    Arena votes
    27,314
    Size
    27B
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1450
    Math
    1429
    Creative writing
    1358
    Instruction following
    1401
    Hard prompts
    1427
    Long queries
    1423
    Multi-turn chat
    1416

    RAM per quant (real file sizes, via unsloth)

    UD-IQ2_XXS 9GB UD-IQ2_M 10GB UD-Q2_K_XL 11GB Q3_K_M 14GB UD-Q3_K_XL 14GB IQ4_XS 15GB Q4_K_M 17GB UD-Q4_K_XL 18GB Q5_K_M 20GB UD-Q5_K_XL 20GB Q6_K 22GB Q8_0 29GB
  42. 42 qwen3-235b-a22b-no-thinking Alibaba · Apache 2.0 98GB est
    Chat Elo (human votes)
    1403(1399-1408)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    38,174
    Size
    235B (22B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    128GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1446
    Math
    1394
    Creative writing
    1368
    Instruction following
    1381
    Hard prompts
    1421
    Long queries
    1411
    Multi-turn chat
    1411

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~133 GB at Q4, ~98 GB at an aggressive dynamic quant.

  43. 43 qwen3-next-80b-a3b-instruct Alibaba · Apache 2.0 23GB
    Chat Elo (human votes)
    1401(1396-1406)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    22,859
    Size
    80B (3B active)
    Context
    262K
    Multimodal
    text only
    Smallest rig that fits
    24GB GPU

    Benchmarks by category (LMArena Elo)

    Coding
    1445
    Math
    1418
    Creative writing
    1316
    Instruction following
    1379
    Hard prompts
    1421
    Long queries
    1390
    Multi-turn chat
    1404

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 23GB UD-IQ1_M 24GB UD-IQ2_XXS 26GB Q2_K 29GB UD-Q2_K_XL 30GB Q3_K_M 38GB UD-Q3_K_XL 36GB IQ4_XS 43GB Q4_K_M 49GB UD-Q4_K_XL 46GB Q5_K_M 57GB UD-Q5_K_XL 57GB Q6_K 66GB Q8_0 85GB
  44. 44 longcat-flash-chat meituan · MIT ?
    Chat Elo (human votes)
    1401(1395-1408)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    11,387
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1474
    Math
    1416
    Creative writing
    1330
    Instruction following
    1389
    Hard prompts
    1426
    Long queries
    1385
    Multi-turn chat
    1393

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  45. 45 qwen3-235b-a22b-thinking-2507 Alibaba · Apache 2.0 86GB
    Chat Elo (human votes)
    1399(1393-1406)
    AA Intelligence (benchmark)
    19.6 · code 22.1
    WebDev Elo (coding)
    n/a
    Arena votes
    8,985
    Size
    235B (22B active)
    Context
    262K
    Multimodal
    text only
    Smallest rig that fits
    96GB

    Benchmarks by category (LMArena Elo)

    Coding
    1442
    Math
    1398
    Creative writing
    1374
    Instruction following
    1386
    Hard prompts
    1418
    Long queries
    1401
    Multi-turn chat
    1393

    RAM per quant (real file sizes, via unsloth)

    Q2_K 86GB UD-Q2_K_XL 88GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 126GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 167GB Q6_K 193GB Q8_0 250GB
  46. 46 deepseek-r1 DeepSeek · MIT ?
    Chat Elo (human votes)
    1398(1393-1403)
    AA Intelligence (benchmark)
    18.5 · code 24.6
    WebDev Elo (coding)
    n/a
    Arena votes
    18,524
    Size
    unconfirmed
    Context
    164K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1445
    Math
    1412
    Creative writing
    1374
    Instruction following
    1397
    Hard prompts
    1419
    Long queries
    1399
    Multi-turn chat
    1410

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  47. 47 qwen3.5-35b-a3b Alibaba · Apache 2.0 11GB
    Chat Elo (human votes)
    1396(1391-1400)
    AA Intelligence (benchmark)
    24.0 · code 37.0
    WebDev Elo (coding)
    1251
    Arena votes
    29,158
    Size
    35B (3B active)
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1435
    Math
    1399
    Creative writing
    1344
    Instruction following
    1388
    Hard prompts
    1413
    Long queries
    1402
    Multi-turn chat
    1394

    RAM per quant (real file sizes, via unsloth)

    UD-IQ2_XXS 11GB UD-IQ2_M 11GB UD-Q2_K_XL 12GB Q3_K_M 16GB UD-Q3_K_XL 17GB Q4_K_M 22GB UD-Q4_K_XL 22GB Q5_K_M 26GB UD-Q5_K_XL 26GB Q6_K 29GB Q8_0 37GB
  48. 48 deepseek-v3-0324 DeepSeek · MIT 186GB
    Chat Elo (human votes)
    1396(1392-1399)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    45,480
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1429
    Math
    1370
    Creative writing
    1390
    Instruction following
    1379
    Hard prompts
    1409
    Long queries
    1393
    Multi-turn chat
    1410

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 186GB UD-IQ1_M 196GB UD-IQ2_XXS 219GB Q2_K 244GB UD-Q2_K_XL 248GB Q3_K_M 319GB UD-Q3_K_XL 321GB Q4_K_M 404GB UD-Q4_K_XL 405GB Q5_K_M 475GB Q6_K 551GB Q8_0 713GB
  49. 49 qwen3-vl-235b-a22b-thinking Alibaba · Apache 2.0 63GB
    Chat Elo (human votes)
    1395(1388-1402)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    7,941
    Size
    235B (22B active)
    Context
    131K
    Multimodal
    yes (images)
    Smallest rig that fits
    64GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1455
    Math
    1405
    Creative writing
    1338
    Instruction following
    1383
    Hard prompts
    1418
    Long queries
    1406
    Multi-turn chat
    1386

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 63GB UD-IQ1_M 70GB Q2_K 86GB UD-Q2_K_XL 88GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 125GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 167GB Q6_K 193GB Q8_0 250GB
  50. 50 step-3.5-flash stepfun · Apache 2.0 ?
    Chat Elo (human votes)
    1395(1391-1399)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    55,991
    Size
    unconfirmed
    Context
    262K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1450
    Math
    1408
    Creative writing
    1346
    Instruction following
    1387
    Hard prompts
    1413
    Long queries
    1406
    Multi-turn chat
    1397

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  51. 51 mimo-v2-flash (non-thinking) Xiaomi · MIT ?
    Chat Elo (human votes)
    1393(1389-1396)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1330
    Arena votes
    46,549
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1446
    Math
    1378
    Creative writing
    1357
    Instruction following
    1384
    Hard prompts
    1414
    Long queries
    1405
    Multi-turn chat
    1389

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  52. 52 minimax-m2.5 minimax · Modified MIT 63GB
    Chat Elo (human votes)
    1390(1386-1394)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1386
    Arena votes
    41,088
    Size
    unconfirmed
    Context
    205K
    Multimodal
    text only
    Smallest rig that fits
    64GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1444
    Math
    1396
    Creative writing
    1358
    Instruction following
    1381
    Hard prompts
    1415
    Long queries
    1403
    Multi-turn chat
    1394

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 63GB UD-IQ1_M 68GB UD-IQ2_XXS 74GB UD-IQ2_M 78GB Q2_K 83GB UD-Q2_K_XL 86GB Q3_K_M 109GB UD-Q3_K_XL 101GB IQ4_XS 122GB Q4_K_M 138GB UD-Q4_K_XL 131GB Q5_K_M 162GB UD-Q5_K_XL 162GB Q6_K 188GB Q8_0 243GB
  53. 53 qwen3-coder-480b-a35b-instruct Alibaba · Apache 2.0 150GB
    Chat Elo (human votes)
    1388(1383-1392)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1272
    Arena votes
    25,694
    Size
    480B (35B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1457
    Math
    1377
    Creative writing
    1366
    Instruction following
    1385
    Hard prompts
    1414
    Long queries
    1409
    Multi-turn chat
    1399

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 150GB Q2_K 175GB UD-Q2_K_XL 180GB Q3_K_M 229GB UD-Q3_K_XL 213GB IQ4_XS 261GB Q4_K_M 290GB UD-Q4_K_XL 276GB Q5_K_M 340GB UD-Q5_K_XL 340GB Q6_K 394GB Q8_0 510GB
  54. 54 mimo-v2-flash (thinking) Xiaomi · MIT ?
    Chat Elo (human votes)
    1387(1381-1393)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1291
    Arena votes
    10,934
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1431
    Math
    1373
    Creative writing
    1336
    Instruction following
    1379
    Hard prompts
    1412
    Long queries
    1402
    Multi-turn chat
    1375

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  55. 55 minimax-m2.1-preview minimax · MIT 63GB
    Chat Elo (human votes)
    1384(1379-1389)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1388
    Arena votes
    17,083
    Size
    unconfirmed
    Context
    205K
    Multimodal
    text only
    Smallest rig that fits
    64GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1439
    Math
    1392
    Creative writing
    1345
    Instruction following
    1385
    Hard prompts
    1407
    Long queries
    1410
    Multi-turn chat
    1391

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 63GB UD-IQ1_M 68GB UD-IQ2_XXS 74GB UD-IQ2_M 78GB Q2_K 83GB UD-Q2_K_XL 86GB Q3_K_M 109GB UD-Q3_K_XL 101GB IQ4_XS 122GB Q4_K_M 138GB UD-Q4_K_XL 131GB Q5_K_M 162GB UD-Q5_K_XL 162GB Q6_K 188GB Q8_0 243GB
  56. 56 qwen3-30b-a3b-instruct-2507 Alibaba · Apache 2.0 9GB
    Chat Elo (human votes)
    1383(1378-1388)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    23,718
    Size
    30B (3B active)
    Context
    262K
    Multimodal
    text only
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1439
    Math
    1380
    Creative writing
    1321
    Instruction following
    1367
    Hard prompts
    1407
    Long queries
    1381
    Multi-turn chat
    1383

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 9GB UD-IQ1_M 10GB UD-IQ2_XXS 10GB UD-IQ2_M 11GB Q2_K 11GB UD-Q2_K_XL 12GB Q3_K_M 15GB UD-Q3_K_XL 14GB IQ4_XS 16GB Q4_K_M 19GB UD-Q4_K_XL 18GB Q5_K_M 22GB UD-Q5_K_XL 22GB Q6_K 25GB Q8_0 32GB
  57. 57 glm-4.6v Zhipu · MIT 37GB
    Chat Elo (human votes)
    1377(1366-1389)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    2,797
    Size
    unconfirmed
    Context
    131K
    Multimodal
    yes (images)
    Smallest rig that fits
    48GB

    Benchmarks by category (LMArena Elo)

    Coding
    1418
    Creative writing
    1344
    Instruction following
    1368
    Hard prompts
    1383
    Long queries
    1378
    Multi-turn chat
    1367

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 37GB UD-IQ1_M 39GB UD-IQ2_XXS 41GB UD-IQ2_M 43GB Q2_K 44GB UD-Q2_K_XL 46GB Q3_K_M 55GB UD-Q3_K_XL 53GB IQ4_XS 58GB Q4_K_M 71GB UD-Q4_K_XL 65GB Q5_K_M 81GB UD-Q5_K_XL 80GB Q6_K 96GB Q8_0 114GB
  58. 58 qwen3-235b-a22b Alibaba · Apache 2.0 86GB
    Chat Elo (human votes)
    1375(1370-1380)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    26,254
    Size
    235B (22B active)
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    96GB

    Benchmarks by category (LMArena Elo)

    Coding
    1433
    Math
    1393
    Creative writing
    1323
    Instruction following
    1357
    Hard prompts
    1392
    Long queries
    1378
    Multi-turn chat
    1372

    RAM per quant (real file sizes, via unsloth)

    Q2_K 86GB UD-Q2_K_XL 88GB Q3_K_M 112GB UD-Q3_K_XL 104GB IQ4_XS 126GB Q4_K_M 142GB UD-Q4_K_XL 134GB Q5_K_M 167GB UD-Q5_K_XL 167GB Q6_K 193GB Q8_0 250GB
  59. 59 glm-4.5-air Zhipu · MIT 39GB
    Chat Elo (human votes)
    1373(1369-1377)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    31,062
    Size
    unconfirmed
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    48GB

    Benchmarks by category (LMArena Elo)

    Coding
    1426
    Math
    1390
    Creative writing
    1328
    Instruction following
    1362
    Hard prompts
    1391
    Long queries
    1379
    Multi-turn chat
    1369

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 39GB UD-IQ1_M 40GB UD-IQ2_XXS 43GB UD-IQ2_M 44GB Q2_K 45GB UD-Q2_K_XL 47GB Q3_K_M 57GB UD-Q3_K_XL 55GB IQ4_XS 60GB Q4_K_M 73GB UD-Q4_K_XL 68GB Q5_K_M 84GB UD-Q5_K_XL 83GB Q6_K 99GB Q8_0 117GB
  60. 60 qwen3-next-80b-a3b-thinking Alibaba · Apache 2.0 23GB
    Chat Elo (human votes)
    1370(1364-1375)
    AA Intelligence (benchmark)
    16.7 · code 17.4
    WebDev Elo (coding)
    n/a
    Arena votes
    13,678
    Size
    80B (3B active)
    Context
    262K
    Multimodal
    text only
    Smallest rig that fits
    24GB GPU

    Benchmarks by category (LMArena Elo)

    Coding
    1421
    Math
    1389
    Creative writing
    1324
    Instruction following
    1359
    Hard prompts
    1384
    Long queries
    1370
    Multi-turn chat
    1350

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 23GB UD-IQ1_M 24GB UD-IQ2_XXS 26GB Q2_K 29GB UD-Q2_K_XL 30GB Q3_K_M 38GB UD-Q3_K_XL 35GB IQ4_XS 43GB Q4_K_M 49GB UD-Q4_K_XL 46GB Q5_K_M 57GB UD-Q5_K_XL 57GB Q6_K 66GB Q8_0 85GB
  61. 61 glm-4.7-flash Zhipu · MIT 9GB
    Chat Elo (human votes)
    1368(1362-1374)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    11,719
    Size
    unconfirmed
    Context
    203K
    Multimodal
    text only
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1423
    Math
    1366
    Creative writing
    1313
    Instruction following
    1351
    Hard prompts
    1387
    Long queries
    1374
    Multi-turn chat
    1363

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 9GB UD-IQ1_M 10GB UD-IQ2_XXS 11GB UD-IQ2_M 11GB Q2_K 11GB UD-Q2_K_XL 12GB Q3_K_M 15GB UD-Q3_K_XL 14GB IQ4_XS 16GB Q4_K_M 18GB UD-Q4_K_XL 18GB Q5_K_M 21GB UD-Q5_K_XL 22GB Q6_K 25GB Q8_0 32GB
  62. 62 gemma-3-27b-it Google · Gemma 7GB
    Chat Elo (human votes)
    1366(1362-1369)
    AA Intelligence (benchmark)
    7.4 · code 10.1
    WebDev Elo (coding)
    n/a
    Arena votes
    47,508
    Size
    27B
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1358
    Math
    1322
    Creative writing
    1348
    Instruction following
    1344
    Hard prompts
    1365
    Long queries
    1364
    Multi-turn chat
    1358

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 7GB UD-IQ1_M 7GB UD-IQ2_XXS 8GB UD-IQ2_M 10GB Q2_K 11GB UD-Q2_K_XL 11GB Q3_K_M 13GB UD-Q3_K_XL 14GB IQ4_XS 15GB Q4_K_M 17GB UD-Q4_K_XL 17GB Q5_K_M 19GB UD-Q5_K_XL 19GB Q6_K 22GB Q8_0 29GB
  63. 63 minimax-m1 minimax · Apache 2.0 ?
    Chat Elo (human votes)
    1364(1359-1368)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    35,163
    Size
    unconfirmed
    Context
    1M
    Multimodal
    text only
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1416
    Math
    1371
    Creative writing
    1318
    Instruction following
    1346
    Hard prompts
    1381
    Long queries
    1365
    Multi-turn chat
    1357

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  64. 64 nvidia-nemotron-3-super-120b-a12b NVIDIA · NVIDIA Open Model 53GB
    Chat Elo (human votes)
    1362(1354-1369)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    7,544
    Size
    120B (12B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    64GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1410
    Math
    1376
    Creative writing
    1304
    Instruction following
    1346
    Hard prompts
    1381
    Long queries
    1361
    Multi-turn chat
    1349

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 53GB UD-IQ2_XXS 53GB UD-IQ2_M 53GB UD-Q2_K_XL 55GB UD-Q3_K_XL 63GB UD-Q4_K_XL 84GB UD-Q5_K_XL 108GB Q8_0 128GB
  65. 65 deepseek-v3 DeepSeek · DeepSeek ?
    Chat Elo (human votes)
    1359(1354-1363)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    21,770
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1388
    Math
    1311
    Creative writing
    1349
    Instruction following
    1344
    Hard prompts
    1350
    Long queries
    1375
    Multi-turn chat
    1374

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  66. 66 mistral-small-2506 mistral · Apache 2.0 ?
    Chat Elo (human votes)
    1358(1352-1363)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    17,698
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1412
    Math
    1339
    Creative writing
    1324
    Instruction following
    1339
    Hard prompts
    1374
    Long queries
    1360
    Multi-turn chat
    1367

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  67. 67 command-a-03-2025 Cohere · CC-BY-NC-4.0 ?
    Chat Elo (human votes)
    1354(1350-1357)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    56,217
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1390
    Math
    1309
    Creative writing
    1336
    Instruction following
    1342
    Hard prompts
    1368
    Long queries
    1369
    Multi-turn chat
    1360

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  68. 68 glm-4.5v Zhipu · MIT 70GB
    Chat Elo (human votes)
    1353(1345-1362)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    4,957
    Size
    unconfirmed
    Context
    66K
    Multimodal
    yes (images)
    Smallest rig that fits
    96GB

    Benchmarks by category (LMArena Elo)

    Coding
    1404
    Math
    1358
    Creative writing
    1310
    Instruction following
    1342
    Hard prompts
    1376
    Long queries
    1341
    Multi-turn chat
    1358

    RAM per quant (real file sizes, via ggml-org)

    Q4_K_M 70GB
  69. 69 gpt-oss-120b OpenAI · Apache 2.0 63GB
    Chat Elo (human votes)
    1352(1348-1357)
    AA Intelligence (benchmark)
    23.8 · code 30.4
    WebDev Elo (coding)
    n/a
    Arena votes
    30,615
    Size
    120B
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    64GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1390
    Math
    1381
    Creative writing
    1278
    Instruction following
    1325
    Hard prompts
    1362
    Long queries
    1324
    Multi-turn chat
    1328

    RAM per quant (real file sizes, via unsloth)

    Q2_K 63GB Q3_K_M 63GB Q4_K_M 63GB UD-Q4_K_XL 63GB Q5_K_M 63GB Q6_K 63GB Q8_0 63GB
  70. 70 step-3 stepfun · Apache 2.0 ?
    Chat Elo (human votes)
    1348(1341-1356)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    6,532
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1408
    Math
    1363
    Creative writing
    1308
    Instruction following
    1345
    Hard prompts
    1378
    Long queries
    1348
    Multi-turn chat
    1342

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  71. 71 llama-3.1-nemotron-ultra-253b-v1 NVIDIA · Nvidia Open Model 105GB est
    Chat Elo (human votes)
    1347(1336-1359)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    2,549
    Size
    253B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    128GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1391
    Math
    1380
    Creative writing
    1332
    Instruction following
    1351
    Hard prompts
    1375
    Long queries
    1341
    Multi-turn chat
    1342

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~143 GB at Q4, ~105 GB at an aggressive dynamic quant.

  72. 72 qwen3-32b Alibaba · Apache 2.0 8GB
    Chat Elo (human votes)
    1347(1338-1357)
    AA Intelligence (benchmark)
    11.5 · code 15.3
    WebDev Elo (coding)
    n/a
    Arena votes
    3,926
    Size
    32B
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1407
    Math
    1399
    Creative writing
    1305
    Instruction following
    1332
    Hard prompts
    1367
    Long queries
    1355
    Multi-turn chat
    1338

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 8GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 12GB Q2_K 12GB UD-Q2_K_XL 13GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 18GB Q4_K_M 20GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 27GB Q8_0 35GB
  73. 73 ling-flash-2.0 ant-group · MIT 36GB
    Chat Elo (human votes)
    1346(1339-1354)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    6,997
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    48GB

    Benchmarks by category (LMArena Elo)

    Coding
    1412
    Math
    1354
    Creative writing
    1269
    Instruction following
    1317
    Hard prompts
    1366
    Long queries
    1326
    Multi-turn chat
    1317

    RAM per quant (real file sizes, via bartowski)

    Q2_K 36GB Q3_K_M 47GB
  74. 74 minimax-m2 minimax · Apache 2.0 64GB
    Chat Elo (human votes)
    1346(1338-1354)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    1296
    Arena votes
    6,862
    Size
    unconfirmed
    Context
    205K
    Multimodal
    text only
    Smallest rig that fits
    64GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1385
    Math
    1355
    Creative writing
    1287
    Instruction following
    1339
    Hard prompts
    1369
    Long queries
    1343
    Multi-turn chat
    1364

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 64GB UD-IQ1_M 69GB UD-IQ2_XXS 74GB UD-IQ2_M 78GB Q2_K 83GB UD-Q2_K_XL 86GB Q3_K_M 109GB UD-Q3_K_XL 101GB IQ4_XS 122GB Q4_K_M 138GB UD-Q4_K_XL 132GB Q5_K_M 162GB UD-Q5_K_XL 162GB Q6_K 188GB Q8_0 243GB
  75. 75 nvidia-llama-3.3-nemotron-super-49b-v1.5 NVIDIA · Nvidia Open 24GB est
    Chat Elo (human votes)
    1343(1333-1353)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    3,344
    Size
    49B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    24GB GPU

    Benchmarks by category (LMArena Elo)

    Coding
    1405
    Math
    1390
    Creative writing
    1308
    Instruction following
    1322
    Hard prompts
    1361
    Long queries
    1343
    Multi-turn chat
    1342

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~31 GB at Q4, ~24 GB at an aggressive dynamic quant.

  76. 76 gemma-3-12b-it Google · Gemma 3GB
    Chat Elo (human votes)
    1342(1332-1351)
    AA Intelligence (benchmark)
    5.5 · code 5.8
    WebDev Elo (coding)
    n/a
    Arena votes
    3,829
    Size
    12B
    Context
    131K
    Multimodal
    yes (images)
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1316
    Math
    1318
    Creative writing
    1334
    Instruction following
    1321
    Hard prompts
    1332
    Long queries
    1346
    Multi-turn chat
    1344

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 3GB UD-IQ1_M 3GB UD-IQ2_XXS 4GB UD-IQ2_M 4GB Q2_K 5GB UD-Q2_K_XL 5GB Q3_K_M 6GB UD-Q3_K_XL 6GB IQ4_XS 7GB Q4_K_M 7GB UD-Q4_K_XL 7GB Q5_K_M 8GB UD-Q5_K_XL 8GB Q6_K 10GB Q8_0 13GB
  77. 77 qwq-32b Alibaba · Apache 2.0 8GB
    Chat Elo (human votes)
    1336(1332-1341)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    25,379
    Size
    32B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1384
    Math
    1364
    Creative writing
    1294
    Instruction following
    1324
    Hard prompts
    1357
    Long queries
    1336
    Multi-turn chat
    1322

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 8GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 11GB Q2_K 12GB UD-Q2_K_XL 13GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 18GB Q4_K_M 20GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 27GB Q8_0 35GB
  78. 78 llama-3.1-405b-instruct-bf16 Meta · Llama 3.1 Community 166GB est
    Chat Elo (human votes)
    1335(1331-1339)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    41,375
    Size
    405B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1375
    Math
    1316
    Creative writing
    1301
    Instruction following
    1313
    Hard prompts
    1341
    Long queries
    1327
    Multi-turn chat
    1339

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~227 GB at Q4, ~166 GB at an aggressive dynamic quant.

  79. 79 llama-3.1-405b-instruct-fp8 Meta · Llama 3.1 Community 166GB est
    Chat Elo (human votes)
    1333(1329-1337)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    59,656
    Size
    405B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1368
    Math
    1319
    Creative writing
    1304
    Instruction following
    1314
    Hard prompts
    1336
    Long queries
    1319
    Multi-turn chat
    1329

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~227 GB at Q4, ~166 GB at an aggressive dynamic quant.

  80. 80 olmo-3.1-32b-instruct Allen AI · Apache 2.0 7GB
    Chat Elo (human votes)
    1330(1324-1336)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    12,211
    Size
    32B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1384
    Math
    1305
    Creative writing
    1291
    Instruction following
    1322
    Hard prompts
    1350
    Long queries
    1337
    Multi-turn chat
    1328

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 7GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 11GB Q2_K 12GB UD-Q2_K_XL 12GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 17GB Q4_K_M 19GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 26GB Q8_0 34GB
  81. 81 molmo-2-8b Allen AI · Apache 2.0 7GB est
    Chat Elo (human votes)
    1328(1307-1350)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    799
    Size
    8B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Instruction following
    1313
    Hard prompts
    1341

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  82. 82 llama-3.3-nemotron-49b-super-v1 NVIDIA · Nvidia 24GB est
    Chat Elo (human votes)
    1328(1316-1340)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    2,218
    Size
    49B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    24GB GPU

    Benchmarks by category (LMArena Elo)

    Coding
    1363
    Creative writing
    1298
    Instruction following
    1326
    Hard prompts
    1361
    Long queries
    1338
    Multi-turn chat
    1331

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~31 GB at Q4, ~24 GB at an aggressive dynamic quant.

  83. 83 qwen3-30b-a3b Alibaba · Apache 2.0 9GB
    Chat Elo (human votes)
    1327(1322-1332)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    26,470
    Size
    30B (3B active)
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1387
    Math
    1352
    Creative writing
    1284
    Instruction following
    1312
    Hard prompts
    1345
    Long queries
    1339
    Multi-turn chat
    1320

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 9GB UD-IQ1_M 10GB UD-IQ2_XXS 10GB UD-IQ2_M 11GB Q2_K 11GB UD-Q2_K_XL 12GB Q3_K_M 15GB UD-Q3_K_XL 14GB IQ4_XS 16GB Q4_K_M 19GB UD-Q4_K_XL 18GB Q5_K_M 22GB UD-Q5_K_XL 22GB Q6_K 25GB Q8_0 32GB
  84. 84 llama-4-maverick-17b-128e-instruct Meta · Llama 4 121GB
    Chat Elo (human votes)
    1327(1323-1331)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    39,950
    Size
    17B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    128GB Mac

    Benchmarks by category (LMArena Elo)

    Coding
    1373
    Math
    1318
    Creative writing
    1307
    Instruction following
    1314
    Hard prompts
    1339
    Long queries
    1335
    Multi-turn chat
    1324

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 121GB UD-IQ1_M 128GB UD-IQ2_XXS 135GB UD-IQ2_M 141GB Q2_K 146GB UD-Q2_K_XL 153GB Q3_K_M 191GB UD-Q3_K_XL 180GB IQ4_XS 214GB Q4_K_M 243GB UD-Q4_K_XL 232GB Q5_K_M 284GB UD-Q5_K_XL 287GB Q6_K 329GB Q8_0 426GB
  85. 85 deepseek-v2.5-1210 DeepSeek · DeepSeek 142GB
    Chat Elo (human votes)
    1323(1315-1332)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    6,795
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1375
    Math
    1292
    Creative writing
    1310
    Instruction following
    1315
    Hard prompts
    1329
    Long queries
    1337
    Multi-turn chat
    1323

    RAM per quant (real file sizes, via lmstudio-community)

    Q4_K_M 142GB Q6_K 194GB Q8_0 251GB
  86. 86 llama-4-scout-17b-16e-instruct Meta · Llama 32GB
    Chat Elo (human votes)
    1323(1318-1327)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    30,265
    Size
    17B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1362
    Math
    1309
    Creative writing
    1290
    Instruction following
    1301
    Hard prompts
    1330
    Long queries
    1327
    Multi-turn chat
    1321

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 32GB UD-IQ1_M 35GB UD-IQ2_XXS 37GB UD-IQ2_M 39GB Q2_K 40GB UD-Q2_K_XL 42GB Q3_K_M 52GB UD-Q3_K_XL 49GB IQ4_XS 58GB Q4_K_M 65GB UD-Q4_K_XL 62GB Q5_K_M 77GB UD-Q5_K_XL 79GB Q6_K 88GB Q8_0 115GB
  87. 87 ring-flash-2.0 ant-group · MIT ?
    Chat Elo (human votes)
    1321(1313-1328)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    7,138
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1391
    Math
    1339
    Creative writing
    1260
    Instruction following
    1317
    Hard prompts
    1350
    Long queries
    1330
    Multi-turn chat
    1283

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  88. 88 llama-3.3-70b-instruct Meta · Llama-3.3 16GB
    Chat Elo (human votes)
    1318(1315-1322)
    AA Intelligence (benchmark)
    9.4 · code 11.9
    WebDev Elo (coding)
    n/a
    Arena votes
    54,723
    Size
    70B
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1345
    Math
    1296
    Creative writing
    1286
    Instruction following
    1292
    Hard prompts
    1321
    Long queries
    1312
    Multi-turn chat
    1317

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 16GB UD-IQ1_M 17GB UD-IQ2_XXS 19GB UD-IQ2_M 24GB Q2_K 26GB UD-Q2_K_XL 27GB Q3_K_M 34GB UD-Q3_K_XL 35GB IQ4_XS 38GB Q4_K_M 43GB UD-Q4_K_XL 43GB Q5_K_M 50GB UD-Q5_K_XL 50GB Q6_K 58GB
  89. 89 gemma-3n-e4b-it Google · Gemma 3GB
    Chat Elo (human votes)
    1318(1313-1323)
    AA Intelligence (benchmark)
    n/a · code 3.2
    WebDev Elo (coding)
    n/a
    Arena votes
    22,565
    Size
    4B
    Context
    33K
    Multimodal
    text only
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1308
    Math
    1260
    Creative writing
    1300
    Instruction following
    1281
    Hard prompts
    1313
    Long queries
    1311
    Multi-turn chat
    1292

    RAM per quant (real file sizes, via unsloth)

    UD-IQ2_XXS 3GB UD-IQ2_M 3GB Q2_K 3GB UD-Q2_K_XL 4GB Q3_K_M 4GB UD-Q3_K_XL 4GB IQ4_XS 4GB Q4_K_M 5GB UD-Q4_K_XL 5GB Q5_K_M 5GB UD-Q5_K_XL 6GB Q6_K 6GB Q8_0 7GB
  90. 90 qwen-max-0919 Alibaba · Qwen ?
    Chat Elo (human votes)
    1318(1312-1324)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    16,478
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1353
    Math
    1291
    Creative writing
    1285
    Instruction following
    1302
    Hard prompts
    1319
    Long queries
    1326
    Multi-turn chat
    1305

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  91. 91 gpt-oss-20b OpenAI · Apache 2.0 11GB
    Chat Elo (human votes)
    1317(1311-1324)
    AA Intelligence (benchmark)
    14.9 · code 20.7
    WebDev Elo (coding)
    n/a
    Arena votes
    10,621
    Size
    20B
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1369
    Math
    1336
    Creative writing
    1238
    Instruction following
    1281
    Hard prompts
    1323
    Long queries
    1301
    Multi-turn chat
    1291

    RAM per quant (real file sizes, via unsloth)

    Q2_K 11GB Q3_K_M 12GB Q4_K_M 12GB UD-Q4_K_XL 12GB Q5_K_M 12GB Q6_K 12GB Q8_0 12GB
  92. 92 nvidia-nemotron-3-nano-30b-a3b-bf16 NVIDIA · NVIDIA Open Model 16GB est
    Chat Elo (human votes)
    1316(1310-1321)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    15,494
    Size
    30B (3B active)
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1363
    Math
    1352
    Creative writing
    1249
    Instruction following
    1293
    Hard prompts
    1328
    Long queries
    1291
    Multi-turn chat
    1299

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~21 GB at Q4, ~16 GB at an aggressive dynamic quant.

  93. 93 mistral-large-2407 mistral · Mistral Research ?
    Chat Elo (human votes)
    1314(1310-1318)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    45,459
    Size
    unconfirmed
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1354
    Math
    1288
    Creative writing
    1287
    Instruction following
    1299
    Hard prompts
    1320
    Long queries
    1304
    Multi-turn chat
    1296

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  94. 94 deepseek-v2.5 DeepSeek · DeepSeek 142GB
    Chat Elo (human votes)
    1307(1302-1312)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    24,572
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1368
    Math
    1288
    Creative writing
    1265
    Instruction following
    1292
    Hard prompts
    1322
    Long queries
    1314
    Multi-turn chat
    1292

    RAM per quant (real file sizes, via lmstudio-community)

    Q4_K_M 142GB Q6_K 194GB Q8_0 251GB
  95. 95 granite-4.1-8b IBM · Apache 2.0 3GB
    Chat Elo (human votes)
    1307(1297-1317)
    AA Intelligence (benchmark)
    n/a · code 9.5
    WebDev Elo (coding)
    1194
    Arena votes
    4,062
    Size
    8B
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1353
    Math
    1321
    Creative writing
    1268
    Instruction following
    1290
    Hard prompts
    1323
    Long queries
    1296
    Multi-turn chat
    1286

    RAM per quant (real file sizes, via unsloth)

    UD-IQ2_M 3GB UD-Q2_K_XL 4GB Q3_K_M 4GB UD-Q3_K_XL 5GB IQ4_XS 5GB Q4_K_M 5GB UD-Q4_K_XL 5GB Q5_K_M 6GB UD-Q5_K_XL 6GB Q6_K 7GB Q8_0 9GB
  96. 96 olmo-3-32b-think Allen AI · Apache 2.0 7GB
    Chat Elo (human votes)
    1306(1297-1314)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    5,933
    Size
    32B
    Context
    66K
    Multimodal
    text only
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1364
    Math
    1311
    Creative writing
    1262
    Instruction following
    1300
    Hard prompts
    1328
    Long queries
    1320
    Multi-turn chat
    1300

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 7GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 11GB Q2_K 12GB UD-Q2_K_XL 12GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 17GB Q4_K_M 19GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 26GB Q8_0 34GB
  97. 97 mistral-large-2411 mistral · MRL ?
    Chat Elo (human votes)
    1305(1301-1310)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    28,073
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1346
    Math
    1282
    Creative writing
    1276
    Instruction following
    1295
    Hard prompts
    1313
    Long queries
    1305
    Multi-turn chat
    1293

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  98. 98 gemma-3-4b-it Google · Gemma 1GB
    Chat Elo (human votes)
    1303(1294-1313)
    AA Intelligence (benchmark)
    n/a · code 2.7
    WebDev Elo (coding)
    n/a
    Arena votes
    4,171
    Size
    4B
    Context
    131K
    Multimodal
    yes (images)
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1274
    Math
    1254
    Creative writing
    1276
    Instruction following
    1268
    Hard prompts
    1284
    Long queries
    1311
    Multi-turn chat
    1272

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 1GB UD-IQ1_M 1GB UD-IQ2_XXS 1GB UD-IQ2_M 2GB Q2_K 2GB UD-Q2_K_XL 2GB Q3_K_M 2GB UD-Q3_K_XL 2GB IQ4_XS 2GB Q4_K_M 2GB UD-Q4_K_XL 3GB Q5_K_M 3GB UD-Q5_K_XL 3GB Q6_K 3GB Q8_0 4GB
  99. 99 mistral-small-3.1-24b-instruct-2503 mistral · Apache 2.0 6GB
    Chat Elo (human votes)
    1303(1299-1308)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    33,189
    Size
    24B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1362
    Math
    1278
    Creative writing
    1271
    Instruction following
    1295
    Hard prompts
    1319
    Long queries
    1322
    Multi-turn chat
    1291

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 6GB UD-IQ1_M 6GB UD-IQ2_XXS 7GB UD-IQ2_M 8GB Q2_K 9GB UD-Q2_K_XL 9GB Q3_K_M 11GB UD-Q3_K_XL 12GB IQ4_XS 13GB Q4_K_M 14GB UD-Q4_K_XL 15GB Q5_K_M 17GB UD-Q5_K_XL 17GB Q6_K 19GB Q8_0 25GB
  100. 100 qwen2.5-72b-instruct Alibaba · Qwen 47GB
    Chat Elo (human votes)
    1303(1299-1307)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    39,406
    Size
    72B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    48GB

    Benchmarks by category (LMArena Elo)

    Coding
    1356
    Math
    1296
    Creative writing
    1254
    Instruction following
    1292
    Hard prompts
    1318
    Long queries
    1317
    Multi-turn chat
    1299

    RAM per quant (real file sizes, via lmstudio-community)

    Q4_K_M 47GB Q6_K 64GB Q8_0 77GB
  101. 101 llama-3.1-nemotron-70b-instruct NVIDIA · Llama 3.1 32GB est
    Chat Elo (human votes)
    1299(1291-1307)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    7,140
    Size
    70B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1328
    Math
    1279
    Creative writing
    1277
    Instruction following
    1280
    Hard prompts
    1310
    Long queries
    1281
    Multi-turn chat
    1293

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.

  102. 102 llama-3.1-70b-instruct Meta · Llama 3.1 Community 32GB est
    Chat Elo (human votes)
    1293(1289-1297)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    55,240
    Size
    70B
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1333
    Math
    1269
    Creative writing
    1257
    Instruction following
    1272
    Hard prompts
    1298
    Long queries
    1294
    Multi-turn chat
    1288

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.

  103. 103 gemma-2-27b-it Google · Gemma license 15GB
    Chat Elo (human votes)
    1289(1286-1292)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    75,754
    Size
    27B
    Context
    8K
    Multimodal
    text only
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1305
    Math
    1246
    Creative writing
    1292
    Instruction following
    1271
    Hard prompts
    1282
    Long queries
    1299
    Multi-turn chat
    1279

    RAM per quant (real file sizes, via lmstudio-community)

    IQ4_XS 15GB Q4_K_M 17GB Q5_K_M 19GB Q6_K 22GB Q8_0 29GB
  104. 104 ibm-granite-h-small IBM · Apache 2.0 ?
    Chat Elo (human votes)
    1287(1279-1296)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    5,682
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1329
    Math
    1279
    Creative writing
    1243
    Instruction following
    1268
    Hard prompts
    1301
    Long queries
    1291
    Multi-turn chat
    1282

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  105. 105 llama-3.1-nemotron-51b-instruct NVIDIA · Llama 3.1 24GB est
    Chat Elo (human votes)
    1286(1276-1296)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    3,749
    Size
    51B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    24GB GPU

    Benchmarks by category (LMArena Elo)

    Coding
    1311
    Math
    1271
    Creative writing
    1261
    Instruction following
    1261
    Hard prompts
    1281
    Long queries
    1270
    Multi-turn chat
    1276

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~32 GB at Q4, ~24 GB at an aggressive dynamic quant.

  106. 106 llama-3.1-tulu-3-70b Allen AI · Llama 3.1 26GB
    Chat Elo (human votes)
    1286(1275-1296)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    2,846
    Size
    70B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1308
    Math
    1264
    Creative writing
    1247
    Instruction following
    1272
    Hard prompts
    1273
    Long queries
    1272
    Multi-turn chat
    1278

    RAM per quant (real file sizes, via unsloth)

    Q2_K 26GB Q3_K_M 34GB Q4_K_M 43GB Q5_K_M 50GB
  107. 107 olmo-3.1-32b-think Allen AI · Apache 2.0 7GB
    Chat Elo (human votes)
    1285(1278-1292)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    8,495
    Size
    32B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1338
    Math
    1297
    Creative writing
    1247
    Instruction following
    1278
    Hard prompts
    1305
    Long queries
    1302
    Multi-turn chat
    1266

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 7GB UD-IQ1_M 8GB UD-IQ2_XXS 9GB UD-IQ2_M 11GB Q2_K 12GB UD-Q2_K_XL 12GB Q3_K_M 16GB UD-Q3_K_XL 16GB IQ4_XS 17GB Q4_K_M 19GB UD-Q4_K_XL 20GB Q5_K_M 23GB UD-Q5_K_XL 23GB Q6_K 26GB Q8_0 34GB
  108. 108 nemotron-4-340b-instruct NVIDIA · NVIDIA Open Model 140GB est
    Chat Elo (human votes)
    1277(1271-1282)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    19,659
    Size
    340B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    192GB

    Benchmarks by category (LMArena Elo)

    Coding
    1308
    Math
    1252
    Creative writing
    1238
    Instruction following
    1258
    Hard prompts
    1280
    Long queries
    1288
    Multi-turn chat
    1256

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~191 GB at Q4, ~140 GB at an aggressive dynamic quant.

  109. 109 llama-3-70b-instruct Meta · Llama 3 Community 32GB est
    Chat Elo (human votes)
    1276(1272-1280)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    156,876
    Size
    70B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1306
    Math
    1258
    Creative writing
    1254
    Instruction following
    1258
    Hard prompts
    1279
    Long queries
    1252
    Multi-turn chat
    1272

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.

  110. 110 command-r-plus-08-2024 Cohere · CC-BY-NC-4.0 ?
    Chat Elo (human votes)
    1276(1269-1282)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    9,866
    Size
    unconfirmed
    Context
    128K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1280
    Math
    1231
    Creative writing
    1263
    Instruction following
    1253
    Hard prompts
    1260
    Long queries
    1281
    Multi-turn chat
    1249

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  111. 111 mistral-small-24b-instruct-2501 mistral · Apache 2.0 9GB
    Chat Elo (human votes)
    1274(1268-1280)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    14,681
    Size
    24B
    Context
    33K
    Multimodal
    text only
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1312
    Math
    1262
    Creative writing
    1227
    Instruction following
    1255
    Hard prompts
    1285
    Long queries
    1279
    Multi-turn chat
    1249

    RAM per quant (real file sizes, via unsloth)

    Q2_K 9GB Q3_K_M 11GB Q4_K_M 14GB Q6_K 19GB Q8_0 25GB
  112. 112 qwen2.5-coder-32b-instruct Alibaba · Apache 2.0 12GB
    Chat Elo (human votes)
    1270(1262-1279)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    5,432
    Size
    32B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1342
    Math
    1270
    Creative writing
    1207
    Instruction following
    1264
    Hard prompts
    1303
    Long queries
    1287
    Multi-turn chat
    1251

    RAM per quant (real file sizes, via unsloth)

    Q2_K 12GB Q3_K_M 16GB Q4_K_M 20GB Q5_K_M 23GB Q6_K 27GB Q8_0 35GB
  113. 113 c4ai-aya-expanse-32b Cohere · CC-BY-NC-4.0 17GB est
    Chat Elo (human votes)
    1267(1262-1272)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    27,124
    Size
    32B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    24GB GPU

    Benchmarks by category (LMArena Elo)

    Coding
    1292
    Math
    1233
    Creative writing
    1228
    Instruction following
    1249
    Hard prompts
    1271
    Long queries
    1289
    Multi-turn chat
    1232

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~22 GB at Q4, ~17 GB at an aggressive dynamic quant.

  114. 114 gemma-2-9b-it Google · Gemma license 5GB
    Chat Elo (human votes)
    1266(1263-1270)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    54,611
    Size
    9B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1271
    Math
    1219
    Creative writing
    1258
    Instruction following
    1244
    Hard prompts
    1257
    Long queries
    1265
    Multi-turn chat
    1251

    RAM per quant (real file sizes, via lmstudio-community)

    IQ4_XS 5GB Q4_K_M 6GB Q5_K_M 7GB Q6_K 8GB Q8_0 10GB
  115. 115 deepseek-coder-v2 DeepSeek · DeepSeek License ?
    Chat Elo (human votes)
    1265(1258-1271)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    15,147
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1342
    Math
    1272
    Creative writing
    1204
    Instruction following
    1252
    Hard prompts
    1288
    Long queries
    1287
    Multi-turn chat
    1235

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  116. 116 qwen2-72b-instruct Alibaba · Qianwen LICENSE 30GB
    Chat Elo (human votes)
    1261(1256-1266)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    37,325
    Size
    72B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1296
    Math
    1273
    Creative writing
    1223
    Instruction following
    1241
    Hard prompts
    1273
    Long queries
    1260
    Multi-turn chat
    1243

    RAM per quant (real file sizes, via bartowski)

    Q2_K 30GB Q3_K_M 38GB IQ4_XS 40GB Q4_K_M 47GB
  117. 117 command-r-plus Cohere · CC-BY-NC-4.0 ?
    Chat Elo (human votes)
    1261(1257-1265)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    77,554
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1272
    Math
    1213
    Creative writing
    1236
    Instruction following
    1240
    Hard prompts
    1251
    Long queries
    1261
    Multi-turn chat
    1236

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  118. 118 phi-4 Microsoft · MIT 6GB
    Chat Elo (human votes)
    1256(1251-1261)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    24,126
    Size
    unconfirmed
    Context
    16K
    Multimodal
    text only
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1306
    Math
    1265
    Creative writing
    1210
    Instruction following
    1245
    Hard prompts
    1278
    Long queries
    1266
    Multi-turn chat
    1241

    RAM per quant (real file sizes, via unsloth)

    Q2_K 6GB Q3_K_M 7GB Q4_K_M 9GB Q5_K_M 10GB Q6_K 12GB Q8_0 16GB
  119. 119 olmo-2-0325-32b-instruct Allen AI · Apache-2.0 12GB
    Chat Elo (human votes)
    1251(1241-1262)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    3,334
    Size
    32B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1279
    Math
    1227
    Creative writing
    1224
    Instruction following
    1228
    Hard prompts
    1261
    Long queries
    1243
    Multi-turn chat
    1243

    RAM per quant (real file sizes, via unsloth)

    Q2_K 12GB Q3_K_M 16GB Q4_K_M 19GB Q5_K_M 23GB Q6_K 26GB Q8_0 34GB
  120. 120 command-r-08-2024 Cohere · CC-BY-NC-4.0 ?
    Chat Elo (human votes)
    1250(1243-1256)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    10,140
    Size
    unconfirmed
    Context
    128K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1281
    Math
    1206
    Creative writing
    1209
    Instruction following
    1235
    Hard prompts
    1256
    Long queries
    1262
    Multi-turn chat
    1211

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  121. 121 ministral-8b-2410 mistral · MRL 7GB est
    Chat Elo (human votes)
    1237(1228-1246)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    4,781
    Size
    8B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1274
    Math
    1214
    Creative writing
    1221
    Instruction following
    1211
    Hard prompts
    1251
    Long queries
    1259
    Multi-turn chat
    1207

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  122. 122 qwen1.5-110b-chat Alibaba · Qianwen LICENSE 41GB
    Chat Elo (human votes)
    1234(1228-1239)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    26,195
    Size
    110B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    48GB

    Benchmarks by category (LMArena Elo)

    Coding
    1279
    Math
    1221
    Creative writing
    1193
    Instruction following
    1217
    Hard prompts
    1245
    Long queries
    1228
    Multi-turn chat
    1212

    RAM per quant (real file sizes, via bartowski)

    Q2_K 41GB
  123. 123 qwen1.5-72b-chat Alibaba · Qianwen LICENSE 33GB est
    Chat Elo (human votes)
    1233(1227-1238)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    39,302
    Size
    72B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    48GB

    Benchmarks by category (LMArena Elo)

    Coding
    1274
    Math
    1209
    Creative writing
    1188
    Instruction following
    1211
    Hard prompts
    1239
    Long queries
    1235
    Multi-turn chat
    1217

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~44 GB at Q4, ~33 GB at an aggressive dynamic quant.

  124. 124 mixtral-8x22b-instruct-v0.1 mistral · Apache 2.0 13GB est
    Chat Elo (human votes)
    1229(1224-1233)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    51,416
    Size
    22B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1277
    Math
    1228
    Creative writing
    1190
    Instruction following
    1214
    Hard prompts
    1243
    Long queries
    1217
    Multi-turn chat
    1187

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~16 GB at Q4, ~13 GB at an aggressive dynamic quant.

  125. 125 command-r Cohere · CC-BY-NC-4.0 ?
    Chat Elo (human votes)
    1226(1221-1231)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    54,036
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1243
    Math
    1176
    Creative writing
    1195
    Instruction following
    1198
    Hard prompts
    1214
    Long queries
    1232
    Multi-turn chat
    1196

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  126. 126 llama-3-8b-instruct Meta · Llama 3 Community 7GB est
    Chat Elo (human votes)
    1223(1219-1227)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    104,642
    Size
    8B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1253
    Math
    1193
    Creative writing
    1196
    Instruction following
    1192
    Hard prompts
    1219
    Long queries
    1207
    Multi-turn chat
    1205

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  127. 127 c4ai-aya-expanse-8b Cohere · CC-BY-NC-4.0 7GB est
    Chat Elo (human votes)
    1223(1216-1230)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    9,818
    Size
    8B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1255
    Math
    1201
    Creative writing
    1189
    Instruction following
    1204
    Hard prompts
    1230
    Long queries
    1248
    Multi-turn chat
    1199

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  128. 128 llama-3.1-tulu-3-8b Allen AI · Llama 3.1 3GB
    Chat Elo (human votes)
    1220(1210-1231)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    2,896
    Size
    8B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1253
    Math
    1206
    Creative writing
    1197
    Instruction following
    1208
    Hard prompts
    1223
    Long queries
    1226
    Multi-turn chat
    1178

    RAM per quant (real file sizes, via unsloth)

    Q2_K 3GB Q3_K_M 4GB Q4_K_M 5GB Q5_K_M 6GB Q6_K 7GB Q8_0 9GB
  129. 129 llama-3.1-8b-instruct Meta · Llama 3.1 Community 2GB
    Chat Elo (human votes)
    1211(1207-1215)
    AA Intelligence (benchmark)
    7.6 · code 5.4
    WebDev Elo (coding)
    n/a
    Arena votes
    49,605
    Size
    8B
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1260
    Math
    1189
    Creative writing
    1177
    Instruction following
    1191
    Hard prompts
    1222
    Long queries
    1223
    Multi-turn chat
    1198

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 2GB UD-IQ1_M 2GB UD-IQ2_XXS 3GB UD-IQ2_M 3GB Q2_K 3GB UD-Q2_K_XL 3GB Q3_K_M 4GB UD-Q3_K_XL 4GB IQ4_XS 4GB Q4_K_M 5GB UD-Q4_K_XL 5GB Q5_K_M 6GB UD-Q5_K_XL 6GB Q6_K 7GB
  130. 130 granite-3.1-8b-instruct IBM · Apache 2.0 5GB
    Chat Elo (human votes)
    1208(1197-1219)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    3,090
    Size
    8B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1287
    Math
    1190
    Creative writing
    1170
    Instruction following
    1192
    Hard prompts
    1230
    Long queries
    1231
    Multi-turn chat
    1155

    RAM per quant (real file sizes, via lmstudio-community)

    Q4_K_M 5GB Q6_K 7GB Q8_0 9GB
  131. 131 qwen1.5-32b-chat Alibaba · Qianwen LICENSE 12GB
    Chat Elo (human votes)
    1203(1197-1209)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    21,741
    Size
    32B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1261
    Math
    1201
    Creative writing
    1131
    Instruction following
    1183
    Hard prompts
    1220
    Long queries
    1223
    Multi-turn chat
    1192

    RAM per quant (real file sizes, via lmstudio-community)

    Q2_K 12GB Q3_K_M 16GB IQ4_XS 18GB Q4_K_M 20GB Q5_K_M 23GB Q6_K 27GB Q8_0 35GB
  132. 132 gemma-2-2b-it Google · Gemma license 2GB
    Chat Elo (human votes)
    1200(1196-1204)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    46,616
    Size
    2B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1193
    Math
    1163
    Creative writing
    1183
    Instruction following
    1171
    Hard prompts
    1185
    Long queries
    1186
    Multi-turn chat
    1164

    RAM per quant (real file sizes, via lmstudio-community)

    IQ4_XS 2GB Q4_K_M 2GB Q5_K_M 2GB Q6_K 2GB Q8_0 3GB
  133. 133 phi-3-medium-4k-instruct Microsoft · MIT 5GB
    Chat Elo (human votes)
    1197(1192-1202)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    25,055
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1230
    Math
    1216
    Creative writing
    1149
    Instruction following
    1176
    Hard prompts
    1212
    Long queries
    1190
    Multi-turn chat
    1139

    RAM per quant (real file sizes, via bartowski)

    Q2_K 5GB Q3_K_M 7GB IQ4_XS 7GB Q4_K_M 9GB Q5_K_M 10GB Q6_K 11GB Q8_0 15GB
  134. 134 mixtral-8x7b-instruct-v0.1 mistral · Apache 2.0 7GB est
    Chat Elo (human votes)
    1196(1192-1201)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    73,503
    Size
    7B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1239
    Math
    1191
    Creative writing
    1159
    Instruction following
    1179
    Hard prompts
    1210
    Long queries
    1181
    Multi-turn chat
    1166

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  135. 135 qwen1.5-14b-chat Alibaba · Qianwen LICENSE 10GB est
    Chat Elo (human votes)
    1190(1183-1197)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    17,839
    Size
    14B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1239
    Math
    1167
    Creative writing
    1137
    Instruction following
    1167
    Hard prompts
    1199
    Long queries
    1191
    Multi-turn chat
    1164

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~12 GB at Q4, ~10 GB at an aggressive dynamic quant.

  136. 136 wizardlm-70b Microsoft · Llama 2 Community 32GB est
    Chat Elo (human votes)
    1184(1175-1193)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    8,214
    Size
    70B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1193
    Math
    1157
    Creative writing
    1198
    Instruction following
    1162
    Hard prompts
    1172
    Long queries
    1181
    Multi-turn chat
    1165

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.

  137. 137 deepseek-llm-67b-chat DeepSeek · DeepSeek License 31GB est
    Chat Elo (human votes)
    1184(1172-1195)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    4,932
    Size
    67B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1217
    Math
    1156
    Creative writing
    1135
    Instruction following
    1156
    Hard prompts
    1174
    Long queries
    1183
    Multi-turn chat
    1151

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~41 GB at Q4, ~31 GB at an aggressive dynamic quant.

  138. 138 granite-3.0-8b-instruct IBM · Apache 2.0 5GB
    Chat Elo (human votes)
    1182(1173-1191)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    6,638
    Size
    8B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1240
    Math
    1197
    Creative writing
    1136
    Instruction following
    1171
    Hard prompts
    1202
    Long queries
    1202
    Multi-turn chat
    1136

    RAM per quant (real file sizes, via lmstudio-community)

    Q4_K_M 5GB Q6_K 7GB Q8_0 9GB
  139. 139 gemma-1.1-7b-it Google · Gemma license 3GB
    Chat Elo (human votes)
    1182(1176-1188)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    23,893
    Size
    7B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1217
    Math
    1161
    Creative writing
    1149
    Instruction following
    1155
    Hard prompts
    1193
    Long queries
    1166
    Multi-turn chat
    1124

    RAM per quant (real file sizes, via bartowski)

    Q2_K 3GB Q3_K_M 4GB IQ4_XS 5GB Q4_K_M 5GB Q5_K_M 6GB Q6_K 7GB Q8_0 9GB
  140. 140 granite-3.1-2b-instruct IBM · Apache 2.0 2GB
    Chat Elo (human votes)
    1178(1167-1190)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    3,188
    Size
    2B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1248
    Math
    1198
    Creative writing
    1147
    Instruction following
    1172
    Hard prompts
    1219
    Long queries
    1222
    Multi-turn chat
    1142

    RAM per quant (real file sizes, via lmstudio-community)

    Q4_K_M 2GB Q6_K 2GB Q8_0 3GB
  141. 141 phi-3-small-8k-instruct Microsoft · MIT ?
    Chat Elo (human votes)
    1170(1164-1176)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    17,766
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1204
    Math
    1193
    Creative writing
    1133
    Instruction following
    1153
    Hard prompts
    1188
    Long queries
    1161
    Multi-turn chat
    1118

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  142. 142 llama-2-70b-chat Meta · Llama 2 Community 32GB est
    Chat Elo (human votes)
    1170(1165-1176)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    38,492
    Size
    70B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1178
    Math
    1137
    Creative writing
    1111
    Instruction following
    1134
    Hard prompts
    1159
    Long queries
    1137
    Multi-turn chat
    1133

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.

  143. 143 llama-3.2-3b-instruct Meta · Llama 3.2 1GB
    Chat Elo (human votes)
    1166(1159-1174)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    7,936
    Size
    3B
    Context
    131K
    Multimodal
    text only
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1176
    Math
    1165
    Creative writing
    1144
    Instruction following
    1145
    Hard prompts
    1167
    Long queries
    1158
    Multi-turn chat
    1148

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 1GB UD-IQ1_M 1GB UD-IQ2_XXS 1GB UD-IQ2_M 1GB Q2_K 1GB UD-Q2_K_XL 1GB Q3_K_M 2GB UD-Q3_K_XL 2GB IQ4_XS 2GB Q4_K_M 2GB UD-Q4_K_XL 2GB Q5_K_M 2GB UD-Q5_K_XL 2GB Q6_K 3GB Q8_0 3GB
  144. 144 granite-3.0-2b-instruct IBM · Apache 2.0 2GB
    Chat Elo (human votes)
    1156(1147-1164)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    6,837
    Size
    2B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1209
    Math
    1169
    Creative writing
    1099
    Instruction following
    1131
    Hard prompts
    1176
    Long queries
    1147
    Multi-turn chat
    1115

    RAM per quant (real file sizes, via lmstudio-community)

    Q4_K_M 2GB Q6_K 2GB Q8_0 3GB
  145. 145 qwq-32b-preview Alibaba · Apache 2.0 12GB
    Chat Elo (human votes)
    1155(1143-1166)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    3,231
    Size
    32B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1173
    Math
    1211
    Creative writing
    1105
    Instruction following
    1146
    Hard prompts
    1168
    Long queries
    1178
    Multi-turn chat
    1144

    RAM per quant (real file sizes, via unsloth)

    Q2_K 12GB Q3_K_M 16GB Q4_K_M 20GB Q5_K_M 23GB Q6_K 27GB Q8_0 35GB
  146. 146 llama2-70b-steerlm-chat NVIDIA · Llama 2 Community 32GB est
    Chat Elo (human votes)
    1154(1141-1167)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    3,585
    Size
    70B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Coding
    1145
    Math
    1114
    Creative writing
    1119
    Instruction following
    1123
    Hard prompts
    1145
    Long queries
    1056
    Multi-turn chat
    1111

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.

  147. 147 mistral-7b-instruct-v0.2 mistral · Apache-2.0 7GB est
    Chat Elo (human votes)
    1149(1142-1155)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    19,402
    Size
    7B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1185
    Math
    1127
    Creative writing
    1104
    Instruction following
    1122
    Hard prompts
    1155
    Long queries
    1133
    Multi-turn chat
    1112

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  148. 148 wizardlm-13b Microsoft · Llama 2 Community 9GB est
    Chat Elo (human votes)
    1149(1139-1158)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    7,044
    Size
    13B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1151
    Math
    1064
    Creative writing
    1140
    Instruction following
    1123
    Hard prompts
    1114
    Long queries
    1149
    Multi-turn chat
    1109

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~11 GB at Q4, ~9 GB at an aggressive dynamic quant.

  149. 149 qwen1.5-7b-chat Alibaba · Qianwen LICENSE 3GB
    Chat Elo (human votes)
    1143(1133-1153)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    4,737
    Size
    7B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1209
    Math
    1121
    Creative writing
    1081
    Instruction following
    1123
    Hard prompts
    1150
    Long queries
    1163
    Multi-turn chat
    1112

    RAM per quant (real file sizes, via bartowski)

    Q2_K 3GB Q3_K_M 4GB Q4_K_M 5GB Q5_K_M 6GB Q6_K 6GB Q8_0 8GB
  150. 150 phi-3-mini-4k-instruct-june-2024 Microsoft · MIT ?
    Chat Elo (human votes)
    1142(1136-1149)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    12,297
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1196
    Math
    1193
    Creative writing
    1096
    Instruction following
    1123
    Hard prompts
    1174
    Long queries
    1105
    Multi-turn chat
    1098

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  151. 151 llama-2-13b-chat Meta · Llama 2 Community 9GB est
    Chat Elo (human votes)
    1141(1134-1148)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    19,174
    Size
    13B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1162
    Math
    1110
    Creative writing
    1091
    Instruction following
    1109
    Hard prompts
    1138
    Long queries
    1139
    Multi-turn chat
    1099

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~11 GB at Q4, ~9 GB at an aggressive dynamic quant.

  152. 152 qwen-14b-chat Alibaba · Qianwen LICENSE 10GB est
    Chat Elo (human votes)
    1138(1127-1149)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    4,964
    Size
    14B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    1197
    Math
    1126
    Creative writing
    1102
    Instruction following
    1117
    Hard prompts
    1141
    Long queries
    1126
    Multi-turn chat
    1095

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~12 GB at Q4, ~10 GB at an aggressive dynamic quant.

  153. 153 gemma-7b-it Google · Gemma license 7GB est
    Chat Elo (human votes)
    1137(1127-1146)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    8,925
    Size
    7B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1167
    Math
    1118
    Creative writing
    1101
    Instruction following
    1103
    Hard prompts
    1153
    Long queries
    1111
    Multi-turn chat
    1041

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  154. 154 codellama-34b-instruct Meta · Llama 2 Community 18GB est
    Chat Elo (human votes)
    1136(1127-1145)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    7,366
    Size
    34B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    24GB GPU

    Benchmarks by category (LMArena Elo)

    Coding
    1159
    Math
    1109
    Creative writing
    1086
    Instruction following
    1103
    Hard prompts
    1132
    Long queries
    1097
    Multi-turn chat
    1073

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~23 GB at Q4, ~18 GB at an aggressive dynamic quant.

  155. 155 phi-3-mini-128k-instruct Microsoft · MIT ?
    Chat Elo (human votes)
    1129(1121-1136)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    20,685
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    unknown

    Benchmarks by category (LMArena Elo)

    Coding
    1154
    Math
    1139
    Creative writing
    1085
    Instruction following
    1099
    Hard prompts
    1129
    Long queries
    1071
    Multi-turn chat
    1057

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  156. 156 phi-3-mini-4k-instruct Microsoft · MIT 1GB
    Chat Elo (human votes)
    1127(1121-1134)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    20,118
    Size
    unconfirmed
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1187
    Math
    1150
    Creative writing
    1073
    Instruction following
    1112
    Hard prompts
    1152
    Long queries
    1107
    Multi-turn chat
    1064

    RAM per quant (real file sizes, via bartowski)

    Q2_K 1GB Q3_K_M 2GB IQ4_XS 2GB Q4_K_M 2GB Q5_K_M 3GB Q6_K 3GB Q8_0 4GB
  157. 157 codellama-70b-instruct Meta · Llama 2 Community 32GB est
    Chat Elo (human votes)
    1118(1100-1137)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    1,143
    Size
    70B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    32GB

    Benchmarks by category (LMArena Elo)

    Instruction following
    1096
    Hard prompts
    1159

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~43 GB at Q4, ~32 GB at an aggressive dynamic quant.

  158. 158 gemma-1.1-2b-it Google · Gemma license 1GB
    Chat Elo (human votes)
    1116(1108-1123)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    10,854
    Size
    2B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1171
    Math
    1109
    Creative writing
    1093
    Instruction following
    1099
    Hard prompts
    1138
    Long queries
    1116
    Multi-turn chat
    1050

    RAM per quant (real file sizes, via lmstudio-community)

    Q2_K 1GB Q3_K_M 1GB IQ4_XS 2GB Q4_K_M 2GB Q5_K_M 2GB Q6_K 2GB Q8_0 3GB
  159. 159 llama-3.2-1b-instruct Meta · Llama 3.2 1GB
    Chat Elo (human votes)
    1111(1103-1118)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    8,045
    Size
    1B
    Context
    60K
    Multimodal
    text only
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1148
    Math
    1124
    Creative writing
    1082
    Instruction following
    1085
    Hard prompts
    1113
    Long queries
    1101
    Multi-turn chat
    1075

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 1GB UD-IQ1_M 1GB UD-IQ2_XXS 1GB UD-IQ2_M 1GB Q2_K 1GB UD-Q2_K_XL 1GB Q3_K_M 1GB UD-Q3_K_XL 1GB IQ4_XS 1GB Q4_K_M 1GB UD-Q4_K_XL 1GB Q5_K_M 1GB UD-Q5_K_XL 1GB Q6_K 1GB Q8_0 1GB
  160. 160 mistral-7b-instruct mistral · Apache 2.0 7GB est
    Chat Elo (human votes)
    1109(1100-1119)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    8,977
    Size
    7B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1144
    Math
    1082
    Creative writing
    1091
    Instruction following
    1085
    Hard prompts
    1113
    Long queries
    1097
    Multi-turn chat
    1079

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  161. 161 llama-2-7b-chat Meta · Llama 2 Community 7GB est
    Chat Elo (human votes)
    1107(1100-1114)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    14,148
    Size
    7B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1102
    Math
    1086
    Creative writing
    1072
    Instruction following
    1069
    Hard prompts
    1096
    Long queries
    1075
    Multi-turn chat
    1076

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  162. 162 gemma-2b-it Google · Gemma license 5GB est
    Chat Elo (human votes)
    1093(1081-1104)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    4,780
    Size
    2B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1136
    Math
    1071
    Creative writing
    1078
    Instruction following
    1069
    Hard prompts
    1111
    Long queries
    1083
    Multi-turn chat
    1034

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~5 GB at Q4, ~5 GB at an aggressive dynamic quant.

  163. 163 qwen1.5-4b-chat Alibaba · Qianwen LICENSE 6GB est
    Chat Elo (human votes)
    1090(1081-1099)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    7,597
    Size
    4B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1131
    Math
    1086
    Creative writing
    1049
    Instruction following
    1068
    Hard prompts
    1094
    Long queries
    1077
    Multi-turn chat
    1054

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~6 GB at Q4, ~6 GB at an aggressive dynamic quant.

  164. 164 olmo-7b-instruct Allen AI · Apache-2.0 7GB est
    Chat Elo (human votes)
    1073(1062-1084)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    6,328
    Size
    7B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    8GB

    Benchmarks by category (LMArena Elo)

    Coding
    1106
    Math
    1054
    Creative writing
    1002
    Instruction following
    1029
    Hard prompts
    1067
    Multi-turn chat
    1040

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~8 GB at Q4, ~7 GB at an aggressive dynamic quant.

  165. 165 llama-13b Meta · Non-commercial 9GB est
    Chat Elo (human votes)
    973(958-989)
    AA Intelligence (benchmark)
    n/a
    WebDev Elo (coding)
    n/a
    Arena votes
    2,391
    Size
    13B
    Context
    n/a
    Multimodal
    n/a
    Smallest rig that fits
    16GB

    Benchmarks by category (LMArena Elo)

    Coding
    882
    Math
    921
    Creative writing
    934
    Instruction following
    916
    Hard prompts
    917
    Multi-turn chat
    891

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~11 GB at Q4, ~9 GB at an aggressive dynamic quant.

  166. kimi-k3 Moonshot AI · open weights ?
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    57.1 · code 76.2
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    unconfirmed
    Context
    1M
    Multimodal
    yes (images)
    Smallest rig that fits
    unknown

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  167. glm-5.2 Zhipu · open weights 217GB
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    51.1 · code 68.8
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    744B (40B active)
    Context
    1M
    Multimodal
    text only
    Smallest rig that fits
    256GB

    Zhipu's newest, and the strongest open model out right now on benchmarks. A 744B mixture-of-experts with only 40B active, so it punches like a giant but runs lighter than its size. Needs a big-RAM machine, but an Unsloth 2-bit dynamic quant squeezes it into about 245GB, so a 256GB Mac can run it. The one to watch. Best for: Top-end local reasoning, coding, agents.

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_S 217GB UD-IQ1_M 228GB UD-IQ2_XXS 238GB UD-IQ2_M 239GB UD-Q2_K_XL 254GB UD-Q3_K_XL 343GB UD-Q4_K_XL 467GB UD-Q5_K_XL 562GB Q8_0 801GB
  168. kimi-k2.7-code Moonshot AI · open weights 304GB
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    41.9 · code 60.8
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    unconfirmed
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    512GB Mac Studio

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 304GB UD-IQ2_XXS 318GB UD-IQ2_M 318GB UD-Q2_K_XL 339GB UD-Q3_K_XL 464GB UD-Q4_K_XL 584GB
  169. hy3-preview Tencent · open weights ?
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    41.2 · code 58.8
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    unconfirmed
    Context
    262K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  170. nex-n2-pro nex-agi · open weights ?
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    41.0 · code 59.1
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    unconfirmed
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    unknown

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

  171. nemotron-3-ultra-550b-a55b NVIDIA · open weights 224GB est
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    37.8 · code 49.3
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    550B (55B active)
    Context
    512K
    Multimodal
    text only
    Smallest rig that fits
    256GB

    RAM per quant

    No trusted GGUF build found yet. Estimated from size: ~307 GB at Q4, ~224 GB at an aggressive dynamic quant.

  172. qwen3.6-27b Alibaba · open weights 10GB
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    37.1 · code 53.7
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    27B
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    16GB

    RAM per quant (real file sizes, via unsloth)

    UD-IQ2_XXS 10GB UD-IQ2_M 11GB UD-Q2_K_XL 12GB Q3_K_M 14GB UD-Q3_K_XL 15GB IQ4_XS 16GB Q4_K_M 17GB UD-Q4_K_XL 18GB Q5_K_M 20GB UD-Q5_K_XL 20GB Q6_K 23GB Q8_0 29GB
  173. qwen3.6-35b-a3b Alibaba · open weights 10GB
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    31.6 · code 41.9
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    35B (3B active)
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    16GB

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 10GB UD-IQ2_XXS 11GB UD-IQ2_M 12GB UD-Q2_K_XL 12GB UD-Q3_K_XL 17GB UD-Q4_K_XL 22GB UD-Q5_K_XL 27GB Q8_0 37GB
  174. step-3.7-flash stepfun · open weights 57GB
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    30.3 · code 39.6
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    unconfirmed
    Context
    262K
    Multimodal
    yes (images)
    Smallest rig that fits
    64GB Mac

    RAM per quant (real file sizes, via unsloth)

    UD-IQ1_M 57GB UD-IQ2_XXS 62GB UD-IQ2_M 62GB UD-Q2_K_XL 66GB UD-Q3_K_XL 89GB UD-Q4_K_XL 122GB UD-Q5_K_XL 146GB Q8_0 209GB
  175. trinity-large-thinking arcee-ai · open weights ?
    Chat Elo (human votes)
    not yet rated
    AA Intelligence (benchmark)
    18.2 · code 25.8
    WebDev Elo (coding)
    n/a
    Arena votes
    n/a
    Size
    unconfirmed
    Context
    262K
    Multimodal
    text only
    Smallest rig that fits
    unknown

    RAM per quant

    No trusted GGUF build found yet. Size unconfirmed, so no RAM estimate.

Showing the current top of the field. Older, lower-ranked releases are hidden by default.

How we rank

Two rankings from two sources. Human votes: LMArena style-controlled Elo, which adjusts for response length and formatting, from blind human-preference votes. Benchmarks: Artificial Analysis Intelligence (republished via OpenRouter), which scores models within days of release, so it catches models the arena has not voted on yet. RAM is taken from real GGUF file sizes where available, otherwise estimated from parameter count.

  • Two rankings, two sources. Human votes is LMArena style-controlled Elo, from millions of blind human-preference votes: the gold standard, but a new model needs weeks of votes before it gets an Elo. Benchmarks is Artificial Analysis Intelligence (republished by OpenRouter), a test-based score that lands within days of a launch. Toggle between them. The newest models, the ones the arena has not voted on yet, are pulled in from OpenRouter and shown with their benchmark score and a "new" tag.
  • Other arena scores: WebDev Elo is LMArena's web-development arena (the closest open proxy for coding). Category numbers (Coding, Math, Creative writing, and so on) are LMArena's own category-filtered Elo, from the same votes, so you can pick a model for a specific job.
  • Category benchmarks: the per-category numbers (Coding, Math, Creative writing, Instruction following, Hard prompts, Long queries, Multi-turn) are LMArena's own category-filtered Elo for that model, from the same human votes. They tell you what a model is relatively good at; a model can rank mid-table overall but punch above its weight in coding or math. The main ranking always uses the overall number.
  • RAM per quant is real: where a trusted publisher (Unsloth, llama.cpp, LM Studio, bartowski) ships a GGUF build, we read the actual file size of every quant from Hugging Face. File size is, in practice, the RAM you need, plus a few GB of headroom for context.
  • Strict matching: a model only gets quant data when its name matches the GGUF repo exactly. "No trusted GGUF build found yet" means exactly that, not a guess at a lookalike. New builds get picked up automatically by the daily refresh.
  • Estimates are labeled: where no GGUF exists, RAM is estimated from parameter count (~0.55 bytes/param at Q4, ~0.4 dynamic) and marked "est".
  • Context window and modality come from OpenRouter's live model data, attached only on an exact match.

Sources: LMArena (human votes), Artificial Analysis (benchmarks, via OpenRouter), Hugging Face (quant file sizes). Parameter counts and licenses are from the model providers' own cards.

Open-source LLM FAQ

What is the best open-source LLM right now? +

It depends which signal you trust. By human votes (LMArena Elo) the leader as of July 27, 2026 is glm-5.1 from Zhipu.

Why is a brand-new model not in the arena ranking? +

LMArena Elo is computed from thousands of blind human votes, which take days to weeks to accumulate after a model launches. So a model released this week has no Elo yet. We catch those releases from OpenRouter and show them in the "Just released" section and the Benchmarks ranking with their Artificial Analysis score, until the arena catches up.

How much RAM do I need to run an open LLM at home? +

The quant file size is, in practice, the RAM you need, plus a few GB of headroom for context. Where a trusted GGUF build exists we list the real file size of every quant. As a rule of thumb, Q4 (the everyday sweet spot) is roughly half the parameter count in GB.

Where can I download open LLM weights? +

Expand any model on this page. Where a canonical GGUF build exists we link it directly, plus Hugging Face GGUF search, MLX builds for Apple Silicon, and Ollama.

What do the category benchmark numbers mean? +

Each expanded model shows LMArena Elo filtered by task category: coding, math, creative writing, instruction following, hard prompts, long queries, and multi-turn chat. They come from the same blind human votes as the main ranking, just sliced by what the conversation was about. Use them to pick a model for a specific job, e.g. the best coder that fits your RAM rather than the best all-rounder.

Which quant should I download? +

Q4_K_M is the everyday default: most of the quality at about a quarter of the full size. If the model is just over your RAM budget, try an Unsloth dynamic quant (UD-), which holds quality better at smaller sizes. Q8_0 is near-lossless if you have RAM to spare.

How is this list ranked and updated? +

Models are ranked by LMArena style-controlled Elo from millions of blind human votes, filtered to open-weight models only. Rankings, quant file sizes, and metadata refresh automatically every day. Last refresh: July 27, 2026.

Want this kind of clarity for your business?

BlueFort AI is the content side of BlueFort IT. When you're ready to actually deploy AI, openly or privately, that's their day job.