Reference — Models

The right model for the job.

The gap between open-weight and frontier models has nearly closed. The question is no longer "can open-weight compete?" It's "which model fits your workload, your compliance requirements, and your budget."

Quick pick

Start here.

Maximum security & US-made

gpt-oss-120b

Apache 2.0, on-premises, US-made (OpenAI). Runs on a single 80GB GPU, with near-frontier reasoning you can fully audit and own.

Maximum intelligence

Claude Fable 5

Via API. Adaptive reasoning, 1M-token context. Pay per use.

Near-frontier, full control

Nemotron 3 Ultra

Hosted dedicated. A US open-weight family you own outright, with strong multi-step reasoning at scale. (GLM-5.2 leads among open-weight models if origin isn't a constraint.)

Closed models

Frontier models (API-only).

Maximum intelligence, minimum control, priced per token.

ModelProviderBest forContext
Claude Fable 5Anthropic · USDeepest reasoning, complex coding, adaptive thinking1M
GPT-5.6 SolOpenAI · USGeneral purpose, coding, agentic workflows1M (est.)
Gemini 3.1 ProGoogle · USMultimodal, video, long context1M
Grok 4.5xAI · USFast, token-efficient reasoning and coding500K
Muse Spark 1.1Meta · USMultimodal reasoning, tool use, agent orchestration262K
Qwen 3.8 MaxAlibaba · ChinaMultimodal enterprise flagship1M

Claude's newest flagship is Opus 5 (July 2026); Fable 5 is the top-capability tier. Pricing changes monthly, so we quote current rates in every proposal

Trade-off: Maximum intelligence. Your data transits the vendor's infrastructure on every API call. The enterprise agreements we configure prevent training on your data.

Open models

Open-weight models (self-hostable).

Full control, self-hosted or on dedicated infrastructure.

ModelLabParametersLicenseBest for
Gemma 4Google · US2B–31B dense + 26B MoEApache 2.0Multimodal, 140+ languages, efficient edge-to-server
Nemotron 3 UltraNVIDIA · US30B–550B (MoE)NVIDIA Open ModelPlans and runs multi-step tasks; tuned for NVIDIA GPUs
gpt-oss-120bOpenAI · US117B (MoE)Apache 2.0On-prem reasoning and tool-use on a single 80GB GPU
Poolside Laguna S 2.1Poolside · US118B MoE (8B active)OpenMDW-1.1Agentic coding; punches above its size
InklingThinking Machines · US975B MoE (41B active)Apache 2.0Multimodal open base built for customization
Mistral Large 3Mistral · France675B (MoE)Apache 2.0Long context, multilingual, general-purpose
Command A+Cohere · Canada218B MoE (25B active)Apache 2.0Sovereign enterprise RAG with auditable citations
DeepSeek V4DeepSeek · China1.6T (MoE)MITEfficient long-context reasoning; lowest cost floor
GLM-5.2Zhipu · China744B (MoE)MITAgentic coding at very low cost
Kimi K3Moonshot · China2.8T (MoE)Kimi K3 LicenseFrontier agentic reasoning, long context
MiniMax M3MiniMax · China428B (MoE)Community LicenseLong-horizon coding agents; native multimodality
Qwen3-235B-A22BAlibaba · China235B MoE (22B active)Apache 2.0Broad benchmark strength, multilingual (open release; the 3.8 Max flagship is API-only)

MoE (Mixture-of-Experts): the model activates only the parts it needs for each request, giving large capacity at lower running cost.

Trade-off: the best open-weight models now sit within a few percent of frontier on public leaderboards. The gap has narrowed from roughly 8% in early 2024 to low single digits, and it keeps closing, though closed models keep an edge at the very top of reasoning and agentic-coding tasks. Zero data leakage. Full audit trail. You own the model. For most business tasks, open-weight is production-ready.

Origin matters for some buyers. For organizations that prefer to keep to domestic models, there is now a strong US-made open-weight slate: Gemma, Nemotron, gpt-oss, Poolside Laguna and Inkling, with other Western options from Mistral (France) and Cohere (Canada), and the leading Chinese labs alongside. We match the model to your policy, not the other way around.

Legal

Licenses explained.

LicenseCommercial use?Can modify?Must share changes?
MIT (DeepSeek, GLM)YesYesNo
Apache 2.0 (Gemma, gpt-oss, Inkling, Mistral, Command A+, Qwen3)YesYesNo
Custom community (Kimi K3, MiniMax)Yes (attribution / revenue caps apply)YesNo
NVIDIA Open / OpenMDW (Nemotron, Poolside)Yes (with limits)YesNo
Proprietary (Claude, GPT, Gemini, Grok, Muse Spark, Qwen 3.8 Max)Via API onlyNoN/A

For regulated enterprises that need legal certainty, we recommend Apache 2.0 models: no usage caps, no license surprises.

Framework

Intelligence vs. Control vs. Cost.

IntelligenceControl →FRONTIER (API)Claude Fable 5GPT-5.6 SolGemini 3.1 Promost intelligent · least controlOPEN-WEIGHT (self-hosted)Gemma 4Nemotron 3gpt-ossGLM-5.2DeepSeek V4Mistral L3near-frontier · full control

More control = open-weight on your hardware. More intelligence = frontier via API. The best setup is often both: open-weight for volume, frontier for hard problems.

Model landscape last verified July 29, 2026. This changes fast, and new models are released weekly. Talk to us for current recommendations.

Contact

Which model is right for you?

We'll map your use case to the right model, and the right deployment mode.

Get a model recommendation