What We Learned Moving Our Agent Loops from Anthropic to GLM
dennispi
18 points
6 comments
August 18, 2026
Related Discussions
Found 5 related stories in 93.6ms across 8,687 title embeddings via pgvector HNSW
- How GLM built its own inference infrastructure whiteros_e · 385 pts · September 17, 2026 · 59% similar
- One month coding with GLM 5.3 Flash ThibWeb · 135 pts · October 02, 2026 · 58% similar
- GLM-5.3 Artificial Analysis Benchmarks apitman · 114 pts · August 18, 2026 · 52% similar
- GLM-5.3-Flash Philpax · 958 pts · August 26, 2026 · 51% similar
- Ox-Alpha Is GLM? jitbit · 57 pts · August 24, 2026 · 51% similar
Discussion Highlights (3 comments)
dongkeren
It makes sense to take cost per task more important than cost per token, but for long running tasks, are more appropriate indicator could be "cost per accepted outcome". Because for a task, it is hard to say: What is the boundary of the task? Do we count failure and retry as the same task? What if the result "looks" good but rejected by the user? How do we measure the extra human effort for verifying and resuming when using the cheap model? In my opinion, "cost" means more than the direct token usage of a finished task.
maurelius2
Finally an evaluation that go beyond benchmarks. Thanks for sharing details and takeaways. I wonder how you come you opted for GLM; how did the selection process look like?
quinncom
AI;DR > That distinction matters when you're debugging an agent, because "the answers got worse after the model swap" is not one failure, it's five