{"slug":"gpt-oss-120b","name":"gpt-oss-120b","domain":"openai.com","verdict":"As of 2026-07-15, ChatGPT, Claude, Gemini, Grok collectively rank gpt-oss-120b #7 of 13 for open-weight llm. Source: https://modelsagree.com/product/gpt-oss-120b (modelsagree.com, CC BY 4.0).","best_rank":7,"categories":1,"brief":{"category":"best-open-weight-llm","title":"Best open-weight LLM","rank":7,"of":13,"top":"DeepSeek-V4","day":"2026-07-19","why":[{"t":"Apache 2.0 licensing","m":["Claude","ChatGPT"],"q":"Mature open tooling, Apache 2.0 licensing, strong reasoning and function calling"},{"t":"strong reasoning and function calling","m":["Claude","ChatGPT"],"q":"strong reasoning and function calling"},{"t":"single 80GB GPU operation","m":["Claude","ChatGPT"],"q":"single-80GB-GPU operation"},{"t":"dependable customizable production choice","m":["ChatGPT"],"q":"a dependable customizable production choice"}],"gap":[{"t":"coding and agentic workflows","m":["Gemini","ChatGPT","Grok"],"q":"coding, reasoning, and agentic workflows"},{"t":"long-context processing","m":["ChatGPT","Grok"],"q":"long-context processing, agentic execution"},{"t":"unusually strong API value","m":["ChatGPT"],"q":"unusually strong API value"}],"fix":[{"t":"text-only capabilities and older performance ceiling","m":["ChatGPT"],"q":"Its text-only capabilities and older performance ceiling now trail newer open-weight leaders."},{"t":"weaker world knowledge and higher hallucinations","m":["Claude"],"q":"Noticeably weaker world knowledge and higher hallucination rates than same-tier peers"},{"t":"safety-tuned refusals frustrate application domains","m":["Claude"],"q":"its safety-tuned refusals frustrate some application domains"}]},"entries":[{"slug":"best-open-weight-llm","title":"Best open-weight LLM","rank":7,"of":13,"score":3,"appearances":2,"modelRanks":{"ChatGPT":5,"Claude":4},"reason":"Apache 2.0 with the best capability-per-GPU in the open field — MXFP4 quantization lets it run on a single 80GB GPU with strong reasoning and adjustable effort levels, making it the most realistic self-host option for teams that must keep data on their own hardware","reasons":[{"model":"Claude","reason":"Apache 2.0 with the best capability-per-GPU in the open field — MXFP4 quantization lets it run on a single 80GB GPU with strong reasoning and adjustable effort levels, making it the most realistic self-host option for teams that must keep data on their own hardware"},{"model":"ChatGPT","reason":"Mature open tooling, Apache 2.0 licensing, strong reasoning and function calling, and single-80GB-GPU operation make it a dependable customizable production choice."}],"fixes":[{"model":"ChatGPT","fix":"Its text-only capabilities and older performance ceiling now trail newer open-weight leaders."},{"model":"Claude","fix":"Noticeably weaker world knowledge and higher hallucination rates than same-tier peers, and its safety-tuned refusals frustrate some application domains — it is not the pick for knowledge-heavy consumer products"}],"updated":"2026-07-15","api":"https://modelsagree.com/api/v1/best/best-open-weight-llm.json"}],"page":"https://modelsagree.com/product/gpt-oss-120b","check":"https://modelsagree.com/check?q=gpt-oss-120b","updated":"2026-08-10T18:18:45.051Z","attribution":"modelsagree.com, CC BY 4.0"}