ModelsAgree
← All leaderboards

gpt-oss-120b

What ChatGPT, Claude, Gemini & Grok actually say · August 2026

Visit openai.com

The verdict

gpt-oss-120b appears in 1 AI-ranked category.

Positioning brief — for the gpt-oss-120b team

Why the models put gpt-oss-120b at #7 for open-weight llm

  • Apache 2.0 licensing Claude · GPTMature open tooling, Apache 2.0 licensing, strong reasoning and function calling
  • strong reasoning and function calling Claude · GPTstrong reasoning and function calling
  • single 80GB GPU operation Claude · GPTsingle-80GB-GPU operation
  • dependable customizable production choice GPTa dependable customizable production choice

What the models credit DeepSeek-V4 (#1) with — and don’t credit gpt-oss-120b

  • coding and agentic workflows Gemini · GPT · Grokcoding, reasoning, and agentic workflows
  • long-context processing GPT · Groklong-context processing, agentic execution
  • unusually strong API value GPTunusually strong API value

What would move the rank — the models’ fix lines, unified

  • text-only capabilities and older performance ceiling GPTIts text-only capabilities and older performance ceiling now trail newer open-weight leaders.
  • weaker world knowledge and higher hallucinations ClaudeNoticeably weaker world knowledge and higher hallucination rates than same-tier peers
  • safety-tuned refusals frustrate application domains Claudeits safety-tuned refusals frustrate some application domains

Restructured from verbatim model output · nothing invented · every quote machine-verified

#7🧠 Best open-weight LLM2/4 models · updated 2026-07-15
GPT #5Claude #4Gemini Grok

Apache 2.0 with the best capability-per-GPU in the open field — MXFP4 quantization lets it run on a single 80GB GPU with strong reasoning and adjustable effort levels, making it the most realistic self-host option for teams that must keep data on their own hardware

GPT Mature open tooling, Apache 2.0 licensing, strong reasoning and function calling, and single-80GB-GPU operation make it a dependable customizable production choice.

Where gpt-oss-120b falls short, per the models

  • GPT Its text-only capabilities and older performance ceiling now trail newer open-weight leaders.
  • Claude Noticeably weaker world knowledge and higher hallucination rates than same-tier peers, and its safety-tuned refusals frustrate some application domains — it is not the pick for knowledge-heavy consumer products

Top alternatives per the models: DeepSeek-V4 · GLM-5.2 · Qwen3 · Llama 4

Watch gpt-oss-120b

Boards re-poll weekly and the models change their minds. One short email only when gpt-oss-120b's standing moves — a rank change, a rival overtaking, or new reasoning from the models. Nothing otherwise.

Embed your ranking badge

gpt-oss-120b ranks #7 for best open-weight llm by AI-model consensus. Put the badge in your README, docs or site — it updates automatically as the models re-rank.

gpt-oss-120b — ranked #7 for Best open-weight LLM by AI models on ModelsAgree
Markdown (README)
[![gpt-oss-120b — ranked #7 for Best open-weight LLM by AI models on ModelsAgree](https://modelsagree.com/badge/gpt-oss-120b.svg)](https://modelsagree.com/best/best-open-weight-llm?utm_source=badge&utm_medium=embed&utm_campaign=badge-gpt-oss-120b)
HTML
<a href="https://modelsagree.com/best/best-open-weight-llm?utm_source=badge&utm_medium=embed&utm_campaign=badge-gpt-oss-120b"><img src="https://modelsagree.com/badge/gpt-oss-120b.svg" alt="gpt-oss-120b — ranked #7 for Best open-weight LLM by AI models on ModelsAgree" height="28"></a>

Rankings are computed from what the models answer, re-polled on demand · raw reasoning shown verbatim · methodology