BullshitBench measures whether AI models challenge nonsensical prompts instead of confidently answering them
BullshitBench measures whether models detect nonsense, call it out clearly, and avoid confidently continuing with invalid assumptions.
deepseek/deepseek-v4-flash-0731 revision to both published benchmark tracks at none and xhigh reasoning.xhigh scored 0.9212 versus 0.5939 for none; in v2, xhigh scored 1.0700 versus 0.6600 for none.310 response rows and their canonical three-judge aggregate rows with no collection, grading, consensus, refusal, or identity errors.100 new nonsense questions in the v2 set.5 domains: software (40), finance (15), legal (15), medical (15), physics (15).The screenshots below follow the same flow as viewer/index.v2.html, starting with the main chart.
Primary leaderboard-style view showing each model's green/amber/red split. The screenshot uses the viewer's 30-day new-model filter so recent additions remain legible.
2. Domain LandscapeDetection mix by domain to compare overall performance vs each domain at a glance.
3. Detection Rate Over TimeRelease-date trend view focused on Anthropic, OpenAI, Google, and DeepSeek.
4. Do Newer Models Perform Better?All-model scatter by release date vs. green rate.
5. Does Thinking Harder Help?Reasoning scatter (tokens/cost toggle in the viewer) vs. green rate.
6. Model Size and WeightsTotal and active parameter scatter views for models with public size metadata.
Benchmark Scope (v2)100 nonsense prompts total.5 domain groups: software (40), finance (15), legal (15), medical (15), physics (15).13 nonsense techniques (for example: plausible_nonexistent_framework, misapplied_mechanism, nested_nonsense, specificity_trap).3-judge panel aggregation (anthropic/claude-sonnet-4.6, openai/gpt-5.2, google/gemini-3.1-pro-preview) using full panel mode + mean aggregation.192 model/reasoning rows.Clear Pushback: the model clearly rejects the broken premise.Partial Challenge: the model flags issues but still engages the bad premise.Accepted Nonsense: the model treats the nonsense as valid.export OPENROUTER_API_KEY=your_key_here export OPENAI_API_KEY=your_openai_key_here # required only for models routed to OpenAI export OPENAI_PROJECT=proj_xxx # optional: force OpenAI requests to a specific project export OPENAI_ORGANIZATION=org_xxx # optional: force organization context
Provider routing is configured per model via collect.model_providers and
grade.model_providers in config (default is OpenRouter), for example:
{"*":"openrouter","gpt-5.3":"openai"}.
./scripts/run_end_to_end.sh
./scripts/run_end_to_end.sh --config config.v2.json --viewer-output-dir data/v2/latest --with-additional-judges
data/latest):./scripts/run_end_to_end.sh --with-additional-judges
./scripts/run_end_to_end.sh --with-additional-judges --serve --port 8877
Then open http://localhost:8877/viewer/index.v2.html.
Use the Benchmark Version dropdown in the filters panel to switch between published datasets (for example v1 and v2).
data/latest.data/v2/latest.drafts/new-questions.md via scripts/build_questions_v2_from_draft.py.CHANGELOG.md.drafts/new-questions.md.questions.v2.json).docs/TECHNICAL.md.MIT. See LICENSE.
Star History| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | whichllm - поиск лучшей LLM модели под оборудование | 0 | 10 | 08-06-2026 |
| 2 | PyPy: A new benchmark runner for PyPy | 0 | 7.35 | 20-06-2026 |
| 3 | Harness Bench: как оценить агентский harness и выбрать связку с моделью | 0 | 6.68 | 02-07-2026 |
| 4 | turnstone - orchestration for tool-using AI agents | 0 | 24.29 | 19-07-2026 |
| 5 | TileRT - Tile-Based Runtime for Ultra-Low-Latency LLM Inference | 0 | 35 | 28-06-2026 |
| 6 | “We cannot choose to become idiots”: The AI cheating scandal roiling Brown University | 0 | 7 | 10-07-2026 |
| 7 | AI Satire... | 0 | 0 | 09-07-2026 |
| 8 | Popular vs. reliable sources—a blind spot in how LLMs assess information | 0 | 8.23 | 30-07-2026 |
| 9 | A.I. Enshittifies Everything | 0 | 10 | 26-06-2026 |