gemini-2.5-flash lost causality-1. gpt-5.4-mini dropped on spatial-1. gpt-5.4-mini failing spatial-1.
August 15, 2026 — 12:09 AM CT
Drift Alerts
- SCORE_DROP openai/gpt-5.4-mini on spatial-1
- REGRESSION gemini/gemini-2.5-flash on causality-1
Provider Status
- OpenAI Unexpected GPT-5.6 Terra responses
- OpenAI Elevated errors in ChatGPT conversations for Free users
- OpenAI Error while creating custom RBAC roles for Enterprise users
- Anthropic Elevated errors for Claude Fable 5
- Anthropic Service disruption on Claude services
- Anthropic Issues reaching status.claude.com
- Anthropic Service disruption on Claude Code on the web and related services
- Anthropic Elevated errors for Claude Mythos 5, Claude Fable 5, and Claude Sonnet 5
- Anthropic Degraded performance for multiple models
- Anthropic Degraded performance for Claude Opus 5
- Anthropic Degraded performance of multiple models
- Anthropic Elevated errors across many models
- Anthropic Elevated errors on Claude Sonnet 5
- Anthropic Elevated errors on Claude Sonnet 5
- Anthropic Degraded performance on Claude Sonnet 5
- Anthropic Error rates across multiple models
- Anthropic Degraded performance on Claude Sonnet 5
- Anthropic Degraded performance on Claude Opus 4.8
- Anthropic Elevated errors across many models
- Anthropic Elevated errors across all models
- Anthropic Elevated errors on Claude Opus 5 and Haiku 4.5
- Anthropic Elevated errors on Claude Opus 5
- Anthropic Elevated errors on Claude Opus 5
- Anthropic Elevated errors for Opus 5
- Anthropic Elevated errors for Claude Fable 5, Claude Sonnet 5, Claude Haiku 4.5, and other models
- Anthropic Elevated errors for Mythos 5, Fable 5, Opus 5 and Claude Haiku 4.5
- Anthropic Sonnet 4.6 and Sonnet 5 errors elevated
- Anthropic Microsoft Office add-in availability
Scorecard
| Model | ambiguity-1 | causality-1 | code-1 | common-sense-1 | logic-1 | math-1 | spatial-1 |
|---|---|---|---|---|---|---|---|
| anthropic/claude-haiku-4-5 | ✓ (4.33) | ✓ (4.67) | ✓ (4.67) | ✓ (3.33) | ✓ (5) | ✓ (5) | ✓ (5) |
| anthropic/claude-opus-4-6 | ✓ (5) | ✓ (4.67) | ✓ (4.67) | ✓ (4.5) | ✓ (5) | ✓ (5) | ✓ (5) |
| anthropic/claude-sonnet-4-6 | ✓ (4.33) | ✓ (4.67) | ✓ (4.5) | ✓ (3.5) | ✓ (5) | ✓ (5) | ✓ (5) |
| gemini/gemini-2.5-flash | ✓ (4.67) | ✗ (3.17)was ✓ (5) | ✓ (4.67) | ✓ (4.5) | ✓ (5) | ✓ (5) | ✓ (5) |
| gemini/gemini-2.5-pro | ✓ (4.67) | ✓ (4.83) | ✓ (4.67) | ✓ (4.83) | ✓ (5) | ✓ (5) | ✓ (5) |
| ollama/llama3 | — | — | — | — | — | — | — |
| openai/gpt-5.4 | ✓ (4.5) | ✓ (5) | ✓ (4.67) | ✓ (4.33) | ✓ (4.67) | ✓ (4.67) | ✓ (5) |
| openai/gpt-5.4-mini | ✓ (4.5) | ✓ (4.5) | ✓ (4.67) | ✓ (4.33) | ✓ (5) | ✓ (5) | ✗ (2.33)was 3.5 |
Model Status
- → anthropic/claude-haiku-4-5 stable
- → anthropic/claude-opus-4-6 stable
- → anthropic/claude-sonnet-4-6 stable
- ↓ gemini/gemini-2.5-flash down
- → gemini/gemini-2.5-pro stable
- → openai/gpt-5.4 stable
- ↓ openai/gpt-5.4-mini down
Raw Data
- Detail log — full responses and judge verdicts per prompt
- JSON — structured data for programmatic access
- Markdown — plain text report
- responses.json — raw model outputs
- judgments.json — raw judge verdicts
- run.log — debug log
- Agent Skill — how to read and interpret this data
- Methodology — how evaluations work