Reference — Models
The right model for the job.
The gap between open-weight and frontier models has nearly closed. The question is no longer "can open-weight compete?" It's "which model fits your workload, your compliance requirements, and your budget."
Quick pick
Start here.
gpt-oss-120b
Apache 2.0, on-premises, US-made (OpenAI). Runs on a single 80GB GPU, with near-frontier reasoning you can fully audit and own.
Claude Fable 5
Via API. Adaptive reasoning, 1M-token context. Pay per use.
Nemotron 3 Ultra
Hosted dedicated. A US open-weight family you own outright, with strong multi-step reasoning at scale. (GLM-5.2 leads among open-weight models if origin isn't a constraint.)
Closed models
Frontier models (API-only).
Maximum intelligence, minimum control, priced per token.
| Model | Provider | Best for | Context |
|---|---|---|---|
| Claude Fable 5 | Anthropic · US | Deepest reasoning, complex coding, adaptive thinking | 1M |
| GPT-5.6 Sol | OpenAI · US | General purpose, coding, agentic workflows | 1M (est.) |
| Gemini 3.1 Pro | Google · US | Multimodal, video, long context | 1M |
| Grok 4.5 | xAI · US | Fast, token-efficient reasoning and coding | 500K |
| Muse Spark 1.1 | Meta · US | Multimodal reasoning, tool use, agent orchestration | 262K |
| Qwen 3.8 Max | Alibaba · China | Multimodal enterprise flagship | 1M |
Claude's newest flagship is Opus 5 (July 2026); Fable 5 is the top-capability tier. Pricing changes monthly, so we quote current rates in every proposal
Trade-off: Maximum intelligence. Your data transits the vendor's infrastructure on every API call. The enterprise agreements we configure prevent training on your data.
Open models
Open-weight models (self-hostable).
Full control, self-hosted or on dedicated infrastructure.
| Model | Lab | Parameters | License | Best for |
|---|---|---|---|---|
| Gemma 4 | Google · US | 2B–31B dense + 26B MoE | Apache 2.0 | Multimodal, 140+ languages, efficient edge-to-server |
| Nemotron 3 Ultra | NVIDIA · US | 30B–550B (MoE) | NVIDIA Open Model | Plans and runs multi-step tasks; tuned for NVIDIA GPUs |
| gpt-oss-120b | OpenAI · US | 117B (MoE) | Apache 2.0 | On-prem reasoning and tool-use on a single 80GB GPU |
| Poolside Laguna S 2.1 | Poolside · US | 118B MoE (8B active) | OpenMDW-1.1 | Agentic coding; punches above its size |
| Inkling | Thinking Machines · US | 975B MoE (41B active) | Apache 2.0 | Multimodal open base built for customization |
| Mistral Large 3 | Mistral · France | 675B (MoE) | Apache 2.0 | Long context, multilingual, general-purpose |
| Command A+ | Cohere · Canada | 218B MoE (25B active) | Apache 2.0 | Sovereign enterprise RAG with auditable citations |
| DeepSeek V4 | DeepSeek · China | 1.6T (MoE) | MIT | Efficient long-context reasoning; lowest cost floor |
| GLM-5.2 | Zhipu · China | 744B (MoE) | MIT | Agentic coding at very low cost |
| Kimi K3 | Moonshot · China | 2.8T (MoE) | Kimi K3 License | Frontier agentic reasoning, long context |
| MiniMax M3 | MiniMax · China | 428B (MoE) | Community License | Long-horizon coding agents; native multimodality |
| Qwen3-235B-A22B | Alibaba · China | 235B MoE (22B active) | Apache 2.0 | Broad benchmark strength, multilingual (open release; the 3.8 Max flagship is API-only) |
MoE (Mixture-of-Experts): the model activates only the parts it needs for each request, giving large capacity at lower running cost.
Trade-off: the best open-weight models now sit within a few percent of frontier on public leaderboards. The gap has narrowed from roughly 8% in early 2024 to low single digits, and it keeps closing, though closed models keep an edge at the very top of reasoning and agentic-coding tasks. Zero data leakage. Full audit trail. You own the model. For most business tasks, open-weight is production-ready.
Origin matters for some buyers. For organizations that prefer to keep to domestic models, there is now a strong US-made open-weight slate: Gemma, Nemotron, gpt-oss, Poolside Laguna and Inkling, with other Western options from Mistral (France) and Cohere (Canada), and the leading Chinese labs alongside. We match the model to your policy, not the other way around.
Legal
Licenses explained.
| License | Commercial use? | Can modify? | Must share changes? |
|---|---|---|---|
| MIT (DeepSeek, GLM) | Yes | Yes | No |
| Apache 2.0 (Gemma, gpt-oss, Inkling, Mistral, Command A+, Qwen3) | Yes | Yes | No |
| Custom community (Kimi K3, MiniMax) | Yes (attribution / revenue caps apply) | Yes | No |
| NVIDIA Open / OpenMDW (Nemotron, Poolside) | Yes (with limits) | Yes | No |
| Proprietary (Claude, GPT, Gemini, Grok, Muse Spark, Qwen 3.8 Max) | Via API only | No | N/A |
For regulated enterprises that need legal certainty, we recommend Apache 2.0 models: no usage caps, no license surprises.
Framework
Intelligence vs. Control vs. Cost.
More control = open-weight on your hardware. More intelligence = frontier via API. The best setup is often both: open-weight for volume, frontier for hard problems.
Model landscape last verified July 29, 2026. This changes fast, and new models are released weekly. Talk to us for current recommendations.
Contact
Which model is right for you?
We'll map your use case to the right model, and the right deployment mode.
Get a model recommendation →