Clear Ideas Research

AI Capability Index

A Clear Ideas benchmark for complex enterprise reasoning. It measures whether models can decompose ambiguous requests, follow dense instructions, preserve structured state, and return complete outputs that survive automated validation.

Index leader Gemini 3.6 Flash
82.8
Reliability91
Capability80.8
Stateful76.4
Price/perf.87.1
Model coverageFrontier + open
Benchmark scopeEnterprise reasoning
Best price/perf.Gemini 3.6 Flash
Efficient frontierGemini 3.5 Flash Lite
Scoring focusValid output is necessary, but not sufficient.

The index rewards complete, structured responses that preserve dependencies and remain useful after validation, rather than treating any valid object as a strong result.

Core capabilityStateful aggregation is the separator.

Strong results collect intermediate outputs, reuse them across dependent steps, and synthesize useful results without losing the structure of the work.

Model configurationSettings are part of the measured system.

Reasoning level, output discipline, latency, and cost all affect whether a model is practical for repeatable enterprise execution.

Leaderboard

Enterprise reasoning under operational constraints.

The Clear Ideas AI Capability Index compares models on normalized scores for capability, stateful aggregation, reliability, sophistication, speed, and price/performance. It is designed for practical enterprise work where useful output, repair effort, and execution readiness all matter.

Capability Index

Normalized index, higher is better

Intelligence Index

Capability quality without speed or price weighting

Capability vs. Price/Performance

Upper-right models combine strong reasoning with efficient execution

Top Model Shape

Capability, stateful aggregation, reliability, speed, and value profile

Reliability Index

First-shot completion weighted by output completeness and sophistication

Capability shape

Useful outputs need more than valid structure.

The next charts separate three related dimensions: whether a model preserves state, whether it designs with enough useful complexity, and whether the resulting structure is genuinely high quality.

  • Stateful aggregation tracks reuse of intermediate values.
  • Sophistication rewards purposeful workflow depth.
  • Quality vs. structure catches shallow but valid outputs.

Stateful Aggregation Index

Intermediate values reused, collected, and synthesized into final outputs

Sophistication Index

Complexity, step variety, variable use, tags, and output contracts

Quality vs. Structure

Separates valid-looking output from genuinely useful output

Detailed scores

Overall, capability, reliability, sophistication, and price/performance.

ModelProviderTierOverallCapabilityStatefulReliabilitySoph.Price/perf.Speed
Gemini 3.6 Flash GoogleFrontier Efficient 82.880.876.491.382.887.193.2
GPT-5.6 Terra OpenAIFrontier82.785.871.392.992.587.075.8
Gemini 3.5 Flash Lite GoogleFrontier Efficient 82.578.875.190.880.786.998.7
Claude Opus 4.8 AnthropicFrontier Efficient 81.985.377.292.487.175.781.6
GPT-5.6 Luna OpenAIFrontier81.886.775.791.494.486.468.9
GPT-5.6 Sol OpenAIFrontier81.689.278.194.995.486.254.7
Claude Sonnet 5 AnthropicFrontier80.885.377.792.487.077.673.1
GPT-5.5 OpenAIFrontier80.488.878.294.794.669.364.8
Claude Opus 5 AnthropicFrontier79.987.079.593.188.984.954.9
Grok 4.3 SpaceXAIFrontier79.381.574.591.181.881.475.1
GPT-5.3 Codex OpenAIFrontier79.283.677.092.788.378.867.1
Claude Fable 5 AnthropicFrontier78.585.977.792.988.058.974.9
Gemma 4 31B GoogleFrontier Efficient 78.077.775.591.584.477.282.5
GPT-5.4 Mini OpenAIFrontier77.186.578.091.393.479.346.5
Gemma 4 12B GoogleFrontier Efficient 76.374.373.191.080.079.284.2
Grok 4.5 SpaceXAIFrontier76.180.274.191.382.873.768.3
GPT-5.4 OpenAIFrontier75.583.477.093.591.572.649.9
Claude Opus 4.7 AnthropicFrontier74.386.480.193.388.965.341.2
Grok 4.2 SpaceXAIFrontier73.582.177.592.085.574.343.9
Meta Muse Spark 1.2 MetaFrontier70.971.964.477.773.371.871.5
GPT-5.4 Nano OpenAIBelow bar68.171.163.878.074.371.458.6
Claude Haiku 4.5 AnthropicBelow bar64.574.478.972.881.571.443.6
GPT-OSS 120B on Groq OpenAIBelow bar59.259.959.565.369.165.265.1
Claude Sonnet 4.6 AnthropicBelow bar41.252.155.748.958.543.112.0
Command A CohereBelow bar41.049.952.558.056.147.614.2
Gemini 3.1 Pro GoogleBelow bar36.036.231.143.041.038.039.9
Gemini 3.5 Flash GoogleBelow bar30.634.929.737.340.833.822.9

Domain breakdown

Different domains expose different reasoning limits.

Each domain stresses a different capability: grounded extraction, abstract planning, state tracking, aggregation, instruction following, and structured output discipline. Scores are normalized indexes so the shape of each model is easier to compare.

ModelMarket reasoningContract analysisStateful aggregationAction synthesisResearch synthesisFinancial reasoning
Gemini 3.6 Flash9389.688.693.494.389.8
GPT-5.6 Terra93.591.393.38591.791.4
Gemini 3.5 Flash Lite92.491.288.590.894.587.7
Claude Opus 4.889.986.889.489.189.486.8
GPT-5.6 Luna92.890.876.581.993.487.7
GPT-5.6 Sol89.388.886.789.288.983.3
Claude Sonnet 591.887.48789.59380.8
GPT-5.588.486.987.58786.386.7
Claude Opus 589.885.587.586.688.282.7
Grok 4.387.285.1758487.381.4
GPT-5.3 Codex85.584.384.484.787.583.3
Claude Fable 592.886.591.689.790.788.1
Gemma 4 31B88.984.986.890.890.285.8
GPT-5.4 Mini84.279.369.182.483.380.9
Gemma 4 12B91.386.579.690.89277.4
Grok 4.588.7837583.990.388.7
GPT-5.484.476.482.977.887.182.5
Claude Opus 4.790.57875.370.67282
Grok 4.287.280.376.582.678.379
Meta Muse Spark 1.295.287.5088.893.591.6
GPT-5.4 Nano88.385.87.587.879.588.3
Claude Haiku 4.588.674.377.685.371.372.5
GPT-OSS 120B on Groq93.590078.289.980.6
Claude Sonnet 4.682.369.472.98500
Command A77.764.3063.371.564.3
Gemini 3.1 Pro0.387.90088.188.7
Gemini 3.5 Flash085.70088.177.4

Below-bar diagnostics

Models below the bar still show the score gap clearly.

Lower-scoring models remain visible so the frontier bar and score distribution are easy to compare.

OpenAIGPT-5.4 Nano
68.1
AnthropicClaude Haiku 4.5
64.5
OpenAIGPT-OSS 120B on Groq
59.2
AnthropicClaude Sonnet 4.6
41.2
CohereCommand A
41.0
GoogleGemini 3.1 Pro
36.0
GoogleGemini 3.5 Flash
30.6

What the benchmark tests

The benchmark stresses dense enterprise reasoning: extraction, research synthesis, state tracking, constraint following, aggregation, and complete machine-readable output.

Why it is hard

The model has to reason about values that do not exist yet, carry them through multiple dependent steps, respect constraints, and produce an object that survives automated validation.

How scoring works

Scores are normalized indexes. Overall performance combines capability, structural discipline, stateful aggregation, coverage, reliability, sophistication, speed, and price/performance so models can be compared without exposing raw execution mechanics. A score of 100 is a theoretical expert ceiling, not a reward for simply producing a valid output. Models can produce useful partial or repair-assisted work; the scoring reflects the extra reliability cost when the first result needs help.

Whitepaper

Read the technical paper behind the index.

The paper explains the benchmark design, scoring model, statistical appendix, and why stateful aggregation is the key capability separating frontier models from partial performers.

Download PDF Clear Ideas AI Capability Index, April 27, 2026