
gpt-oss-120b
OpenAI · #11 in our ranking · Best on a single GPU
gpt-oss-120b is an open-weight mixture-of-experts model from OpenAI, released Aug 2025, with 117B parameters (5.1B active per token), a 128K-token context window, and the Apache 2.0.
gpt-oss-120b is OpenAI's open-weight reasoning model under Apache 2.0. It activates 5.1B parameters per token and fits on a single 80GB GPU, which makes it a common choice for private, low-cost deployments.
gpt-oss-120b specs
- Provider
- OpenAI
- Architecture
- Mixture-of-experts
- Parameters
- 117B total · 5.1B active
- Context
- 128K tokens | 98K words
- License
- Apache 2.0 Commercial use
- Input
- Text
- Released
- Aug 2025
- Weights
- Hugging Face
Sourced from the model card on Hugging Face. Last checked 2026-10-01.
Highlights
- Apache 2.0 license
- Runs on one 80GB GPU
- Configurable reasoning effort
Where it fits for payment teams
- Transaction monitoring
#3 pick. Apache 2.0 and fits on a single 80GB GPU, a simple way to run triage entirely inside your network.
Harnesses in Runtime
Run gpt-oss-120b through any of these harnesses on Runtime.
How to run gpt-oss-120b
Download the weights
Pull openai/gpt-oss-120b from Hugging Face. Check the Apache 2.0 before you deploy.
Serve it in your cloud
Run it with vLLM or SGLang on your own GPUs, or use a managed provider that hosts the model.
Put it to work in Runtime
Point an agent or a single skill at the model through OpenCode, then compare it to your current model with evals.
gpt-oss-120b FAQ
What is the context window of gpt-oss-120b?
+
gpt-oss-120b supports a 128K-token context window (131,072 tokens), roughly 98K words of English text.
How many parameters does gpt-oss-120b have?
+
gpt-oss-120b is a mixture-of-experts model with 117B total parameters, of which 5.1B are active per token.
Is gpt-oss-120b free for commercial use?
+
Yes. gpt-oss-120b is released under the Apache 2.0 license, which is permissive and allows commercial use and fine-tuning.
What can gpt-oss-120b take as input?
+
gpt-oss-120b accepts text input and generates text.
Is gpt-oss-120b good for transaction monitoring?
+
Yes. gpt-oss-120b is our #3 pick for transaction monitoring. Apache 2.0 and fits on a single 80GB GPU, a simple way to run triage entirely inside your network.
Can I run gpt-oss-120b in my own cloud with Runtime?
+
Yes. Download the weights from Hugging Face (openai/gpt-oss-120b) and serve them with vLLM or SGLang in your cloud, or use a managed provider that hosts the model. Runtime agents can then use gpt-oss-120b through a harness like OpenCode, with your data staying in your environment.