OpenAI
OpenAI: gpt-oss-120b
openai/gpt-oss-120b
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases.
- Context window: 131,072 tokens
- Input: text
- Output: text
- Pricing: $0.037/M input tokens, $0.17/M output tokens
Community sentiment
Negative (-19) — 3 positive, 6 negative, 4 neutral across 13 community samples from X and the web, as of 2026-09-13.
Community discussion of gpt-oss-120b is mixed but leans negative: multiple users report reliability and behavior problems (agent flatlining in tests, frequent incorrect tool calls) and some benchmarks place it behind smaller models. Others highlight strengths such as low resource requirements (running without a discrete GPU) and large context window/spec parity with peers.
- “Most "0% agentic" rows aren't the model dying. They're the oracle choking. openllms B8: five models flatlined at 0% on Docker-less https://t.co/2kjQXzik06 boxes (gemma-4-31b, opus-35b-a3b, glm-4.5-air, gpt-oss-120b, laguna-s-2.1). Agent ran on the host. Completion oracle still” — X
- “Stuck with GPT OSS 120. Lots of tool call failures. Any tips? - Reddit — 2 days ago · GPT OSS 120 is trying to make a lot of incorrect tool calls. GPT OSS is ancient. It can do tool usage but fails. OSS 120B is worst than a modern 9B model for ...For those of you forced to only u” — the web
- “Best Desktops for Local AI in 2026: Lab-Tested Leaderboard — 4 days ago · The Z2 Mini G1a ran GPT-OSS 120B with no discrete GPU at all. HP's mini workstation puts AMD's Ryzen AI Max+ PRO silicon in a compact, quiet, IT-friendly ...” — the web
- “OpenAI models: All the models and what they're best for - Zapier — # OpenAI models: Every model (including GPT-6) and what it's best for ## OpenAI models at a glance | **Model** | **Best for** | **Inputs** | **Outputs** | **Context window** | **Pricing (input / output per M token” — the web
RuntimeWire coverage
- Cerebras details CS-4 rack and targets 10,000 tokens per second with CS-5
- OpenAI reports up to 3.6x lower Jalapeno latency in engineering tests
- OpenAI says Jalapeno beats Nvidia Blackwell on inference speed per watt
- Ox Alpha filters domestic Chinese political risks, CTGT audit finds
- Inco AI releases DFlash 2, reports 21% longer accepted token sequences
- Head to head: DeepSeek-V4-Flash-0731 vs gpt-oss-120b
- Head to head: gpt-oss-120b vs Phi-4-reasoning
- Head to head: gpt-oss-120b vs cohere-command-a
- webAI releases TwiL models to check AI reasoning on consumer hardware
- CTGT says DeepSeek distillation lifted finance scores without censorship transfer
- Head to head: Google: Gemini 3.6 Flash vs gpt-oss-120b
- Head to head: gpt-oss-120b vs Kimi K3
View on OpenRouter. Model data sourced from OpenRouter.