GulliBench: Measuring Skepticism in Frontier Models
rigelbm
18 points
2 comments
August 12, 2026
Related Discussions
Found 5 related stories in 97.0ms across 8,687 title embeddings via pgvector HNSW
- Show HN: JevBench, a reproducible benchmark for typed decision models florianstandhar · 88 pts · September 22, 2026 · 57% similar
- GLM-5.3 Artificial Analysis Benchmarks apitman · 114 pts · August 18, 2026 · 55% similar
- FrontierFinance: The largest open benchmark for investor workflows ashwinpp · 16 pts · July 09, 2026 · 54% similar
- MentalHealthBench gmays · 14 pts · September 24, 2026 · 54% similar
- Terminal-Bench-Science: Evaluating AI agents on scientific research workflows matt_d · 61 pts · August 28, 2026 · 52% similar
Discussion Highlights (2 comments)
dvaplima
I liked approach to evaluating AI behavior, the fact that additional reasoning doesn’t improve performance is quite interesting!
Betaantunes11
interesting approach