GulliBench: Measuring Skepticism in Frontier Models
rigelbm
18 points
2 comments
August 12, 2026
Related Discussions
Found 5 related stories in 43.4ms across 4,128 title embeddings via pgvector HNSW
- GLM-5.3 Artificial Analysis Benchmarks apitman · 114 pts · August 18, 2026 · 55% similar
- FrontierFinance: The largest open benchmark for investor workflows ashwinpp · 16 pts · July 09, 2026 · 54% similar
- Show HN: Reproducibility Benchmark a Risk Quantitative Model mingshi_tz · 11 pts · July 26, 2026 · 50% similar
- Measuring reward-seeking by instilling contrastive beliefs mfiguiere · 11 pts · July 21, 2026 · 48% similar
- The Benchmarkpocalypse cyndunlop · 42 pts · August 18, 2026 · 48% similar
Discussion Highlights (2 comments)
dvaplima
I liked approach to evaluating AI behavior, the fact that additional reasoning doesn’t improve performance is quite interesting!
Betaantunes11
interesting approach