LLMs: Intelligence vs. Cost

theanonymousone 75 points 39 comments September 02, 2026
openteams.com · View on Hacker News

Discussion Highlights (14 comments)

asf1289

Open Teams originally wanted to rent out open source developers to sponsors with Oliphant controlling everything. Now they pivoted to installing local LLMs (on what hardware exactly?). What will happen is that this will be the third consultancy with a lofty narrative after Enthought and Anaconda that Oliphant established. It is always bait-and-switch.

fxwin

> Why AA’s plot is misleading > The first issue I have with it is that it uses a logarithmic scale on the cost axis. Using a log scale is the only way to make you spot the difference between a model that costs $0.015 per task and one that costs $0.032, while the same plot contains a model that costs $3.69 — almost 250 times as expensive. However, the net result is that the viewers can no longer appreciate the immensity of the price difference between the cheap models and the heavy ones; nor can they realize how inconsequential the price differences are between the cheap models. This is an asinine complaint, and nobody can seriously tell me that the last plot on their page [0] is more readable than the AA one [1]. If I'm using a model at the lower range of the cost scale for whatever list of tasks, and i switch to another model at the lower end of the cost scale, my spending might double anyways! This should be reflected in the plot, and linear scale doesn't do it justice. It's also much easier to see the mentioned pareto frontier in the log plot than in the linear one. I can see why they disagree with the pricing determination for open/local models, but I don't think there is one clear right way to do it. So how do they do it instead? >Hardware is priced at zero, on the basis that both an RTX 3090 PC and a 64GB Strix Halo are desirable gaming/work machines anyways. ...oh Would have been nice to mention explicitly how the pareto frontier changes with those new calculations. [0] https://openteams.com/wp-content/uploads/2026/09/all_models-... [1] https://artificialanalysis.ai/#intelligence-comparison-tabs

sinuhe69

When I accessed the site, it showed the FBI badge says that the site was blocked and redirect to fbi.gov !! WTH?

Primer81

to choose a model, you have to consider intelligence, price, and tokens per second. would be nice to see the 3 dimensional plot.

sarjann

Why use electricity alone for open models? Surely you’d want to spread the cost of hardware over the period too?

oliwary

This looks great! I also think speed should be part of the metric (i.e. how long does the model take to actually solve a task). For me, I prefer to run expensive models such as Sol on light reasoning, which usually gives me good answers with quick responses. For my style of coding (quick back-and-forths and corrections) it makes a big difference if a model comes back in 1-2 minutes compared to 5-10, and I am happy to pay a bit extra for that.

vb-8448

Complaining about "bad charting" and posting a chart with y-axis that doesn't start at 0 is kinda weird.

andai

Well done. The inability to switch between log and linear always bothered me. Another thing is if you're using the subscriptions with OpenAI or Anthropic you get an order of magnitude discount relative to the per-token price. So you need to move their models ~10x to the left on the plots to get a fair comparison.

shelled

Is there a case in which a heavy agentic coding user of mid or mid++ tier (remotely hosted) models is better off using PAYG/API pricing than just getting a subscription? (Assuming no easy access to high end local hardware and I've deliberately left the top tier/cutting edge models out becau).

themgt

There is an immense difference in cost between the state-of-the-art models from Anthropic and OpenAI and the much cheaper Chinese models ... How much extra intelligence emptying the wallet purchases obeys the law of diminishing returns: while a top-tier engineer or scientist is probably going to be able to appreciate how much better Fable 5.1 [is] ... most people will have a hard time doing so. Nebari is officially listed as a JATIC product as part of the next-gen toolchain supporting DoD AI development. Are we officially ~one degree of Kevin Bacon from the DoD endorsing running Chinese OSS models because they're self-hosted and we're all too dumb to tell the difference? https://openteams.com/open-source-isnt-the-real-risk-in-nati...

datadrivenangel

The complaint about not being able to switch between linear and log is valid, which is what I did for making a 3D speed/cost/quality frontier application for a recent meetup talk: https://www.williamangel.net/apps/model_performance.html Because speed is important, as the reasoning and hardware determine both cost and speed. it's a three dimensional tradeoff.

esafak

Just give me the logarithmic chart; I'm not an idiot, and it is easier to read. As another reader said, maybe they could make it a toggle.

akazantsev

> This is fine in most cases, but for open-weights models it can be a lot more expensive than what the exact same model can be rented for from third-party API providers. Filter by quantization, and most providers will have the same price. There is some "base" price even for open-weight models. Anything cheaper means some tricks on the provider's side.

SturgeonsLaw

I've been very impressed with GLM 5.3 Flash's performance even without considering cost, but once you factor that in, it's incomparable. Not surprised to see its position on the chart.

Semantic search powered by Rivestack pgvector
5,346 stories · 48,358 chunks indexed