PANONDA
← News
Negócios·3 min read

AT&T says open models cut its AI bill by up to 80%. Here's what that means for your stack.

The New York Times reports enterprises are shifting workloads from closed frontier models to open-weight ones, with AT&T claiming savings of up to 80%. For publishers, the lesson is not "switch models" — it's "stop paying frontier prices for chores."

By PANONDA Newsroom

Foto: Pavel Danilyuk · Pexels

What happened

The New York Times reported this week that open-source AI models are becoming a default option inside large companies, not a hobbyist alternative. The number carrying the piece: AT&T says it is saving as much as 80% on AI costs by moving work away from closed, proprietary models toward open-weight ones. That figure comes from the company, via the Times — not from an audited filing.

Same week, the media side kept moving in the opposite direction on spend. PR Newswire's monthly media recap logged publishers shipping paid AI features, including the New York Post launching its own AI chatbot. So you have enterprises cutting model costs while media companies add AI products on top of their existing bills.

What it means

The misreading here is easy, so let's kill it first. This is not a story about open models being as good as frontier models — it's a story about most AI work not needing a frontier model at all.

A company like AT&T is not routing its hardest reasoning problems to an open model to save money. It's routing the boring volume: ticket summaries, classification, tagging, internal search, transcript cleanup. That work is high-volume and low-difficulty, which is exactly where per-token pricing quietly eats a budget. Move it to a cheaper model and the bill drops a lot without anyone noticing a quality difference.

If you publish for a living, your usage profile looks more like AT&T's than you'd think. Look at where your tokens actually go. It is almost never the writing. It's transcription cleanup, chapter markers, show notes, title variants, repurposing one video into eight posts, tagging an archive, answering "what did I say about X in 2023." Chores. You are probably paying premium rates for chores.

The second implication is negotiating position. Once open-weight models are a credible fallback for routine work, closed-model pricing has a ceiling. That pressure shows up in the tools you buy, not just in raw API pricing. Any vendor charging you a subscription that is mostly a wrapper around someone else's API now has a cost floor that is falling. Expect either cheaper tiers or more generous limits at the same price over the next few quarters. Ask for it.

Third: this cuts against the reflex to buy every AI feature a platform ships. The Post launching a chatbot is a product decision with a revenue theory behind it. Adding an AI subscription to your workflow because it exists is not.

What this does not mean: run models yourself. Self-hosting an open model means GPUs, uptime, and someone who can debug it at 2am. That's not a one-person newsletter operation. The practical version is using open models through a hosted provider, at a lower per-token price, for a specific set of tasks.

Also worth saying plainly: the 80% is a company claim about a company workload. Your savings will be smaller, because your volume is smaller and your fixed subscription costs don't move.

What to do about it

  1. Pull your last three months of AI invoices — API and subscriptions both — and write down the single number you spend per month. Most people don't know it.
  2. Sort your AI tasks into two buckets: judgment (drafting, editing, anything with your name on it) and chores (transcripts, tags, summaries, reformatting).
  3. Move one chore to a cheaper model this week. Transcript cleanup is the easiest test. Run 10 files through both, compare output side by side.
  4. If a tool you pay for is a thin wrapper on an API, email support and ask what their pricing looks like in 90 days. The answer tells you whether to renew annually.
  5. Set a re-check date on your calendar for January. Model pricing is moving fast enough that a stack you optimized in September will be wrong by winter.
  6. Don't self-host. If someone pitches you on running your own model, ask who's on call.

Receba o resumo diário

Todo dia útil, às 7h. 5 minutos de leitura. O que importa, em português, com o ângulo do mercado brasileiro.

Sem spam. Cancele quando quiser.