FM Logo
AI BlogAI NewsAI LabBooksAboutPortfolio
How can I help?
How can I help?

INSIGHT #45SundAI Blog

Where does competitive advantage end when the token costs half and the weights become open?

#Costi AI/#Open source/#Agenti autonomi/#Infrastruttura AI/#Vendor lock-in/#Mistral

AuthorFabrizio Mazzei10/11/20269 min read
Where does competitive advantage end when the token costs half and the weights become open?. AI-generated image

Image generated entirely with AI.

TL;DR

"The price of tokens plunges by 52%, Mistral unveils open weights, and agents enter workflows. The week in which margin shifts from infrastructure to those who use it well."

Loading audio player...
  1. 01Who pays the bill?
  2. 02Do open weights change the game?
  3. 03Do watermarks actually hold up?
  4. 04Do universal agents change anything?
  5. 05What to watch now?

The price of tokens went from $2.07 per million to $0.99 in five months, a 52% drop. In the same week Mistral opened the preview of a 1-trillion-parameter model that runs in European datacenters, with weights coming at the end of the month. Two pieces of news that seem separate and yet tell the same story: the competitive advantage in AI is shifting places.

If a token costs little and a frontier model is downloadable, the moat is no longer the model. It's what you build on top of it. The week offers several concrete proofs of where that advantage is shifting, plus a few accounting traps to keep an eye on.

Where does competitive advantage end when the token costs half and the weights become open?. AI-generated imageImage generated entirely with AI.

Who pays the bill?

Let's start with the numbers that make noise. Meta classified the construction of its AI data centers as "pilot models," tying into a research tax credit introduced in the Reagan era. The result: $2 billion saved in 2024 and nearly $4 billion in fiscal year 2025, according to The New York Times. Meta is the largest known beneficiary among public companies.

The discordant note comes from documents filed by the company itself: it might have to repay what it obtained due to "uncertainties with our research credits." The IRS is already trying to claw back $355 million for another 2013 tax workaround involving News Feed development.

Why it matters in practice. If a slice of the data center bill is covered by research credits, the real price of inference in the list prices is distorted. Anyone building a business model on the subsidy is borrowing margin from the future. The same AI companies are maximizing invested capital with tax leverage as a multiplier, and it's accounting fluff until you see real operating margin.

In the same week, the counterpoint arrives. Silicon Data's index of effective costs went from a peak of $2.07 per million tokens in late May to $0.99 on October 7, a 52% drop. Weekly tokens on OpenRouter grew more than 250-fold between January 2025 and September 2026. The unit price collapses while volumes explode.

The push comes from list prices. On July 30 OpenAI cut API prices for GPT-5.6 Luna by 80% and those for Terra by 20%, bringing Luna's input to $0.20 per million. Anthropic differentiated Claude 5.5 pricing by type of work. The most interesting data point in OpenRouter's analysis is not the discount, but the 32% of customers who kept paying full price after promotions ended: those who return pay because they found real value.

The cost per token falls, but the metric that matters remains cost per completed task.

A cheaper token allows longer reasoning cycles, more verification calls, agents that check each other. The real gain is being able to afford more tokens to reach a correct result. A deep dive on the collapse in inference costs and on the economics of autonomous pipelines helps frame the scale of the phenomenon.

Do open weights change the game?

Mistral Large 4 is the largest model ever built by the French company: 1 trillion parameters with 49 billion active, natively multimodal, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs inside Mistral's European datacenters. The preview is on Mistral Studio, weights arrive at the end of the month.

The number that shifts the work is 82% in the test that asks to reproduce and then patch a real vulnerability, the highest score of any model. Here, though, the caveat is realistic: the barrier is not only technical. You need GPUs, fine-tuning skills, and regulatory oversight that few teams already have in house, and in the projects I follow it is almost always the second bottleneck that makes timelines slip. Before rewriting the architecture, it's worth understanding what you're actually paying for today, with an audit of the AI stack and costs. The interesting detail is that closed models like Claude Opus 5.5 and GPT-6 Astra score almost zero on the same tests because they refuse the task.

Translated for those designing security pipelines: open weights plus self-deployment on private cloud change how malware analysis workflows, vulnerability prioritization, and detection rule writing can be put together. The provider filter stops blocking the pipeline halfway through an incident.

On the open side there is also the NASA + IBM release of the Lunar Foundation Model, trained on nearly 2 million tile bundles derived from 17 years of Lunar Reconnaissance Orbiter data. The move is smart: raw data plus already-trained weights, so small teams skip months of preprocessing and go straight to the problem. The technical question remains one: does it generalize or is it overfit to the LRO dataset?

The picture closes with EmbeddingGemma 2, under 1B parameters, Apache 2.0 license, a single 768-dimensional vector space for text, code, images, video, and audio. For those building retrieval, a multimodal embedding with a permissive license is free leverage. In this scenario, the theme of overcoming vendor lock-in becomes the most useful frame for reading releases week by week.

Do watermarks actually hold up?

OpenAI introduced textGrain, an invisible watermark in texts generated by ChatGPT and Codex for European users. The mechanism works inside the model: during generation it slightly modifies token probabilities and builds a statistical pattern that a dedicated detector can look for. No hidden characters, no invisible spaces.

The release comes just over two months after the transparency obligations of the AI Act. Article 50 calls for machine-readable, effective, interoperable, and robust marking. Access to the detector remains initially reserved for authorized researchers and organizations, so enterprise-side verification remains in the provider's hands.

Tests published by OpenAI show that short texts, translations, and post-generation edits weaken detectability sharply. A watermark that holds up only on raw output is of little use: texts leaving pipelines go through rewriting, summarizing, translation, and formatting. Each pass erodes the signal and makes the detector less reliable.

Google opened SynthID Detector to the public, with over 180 billion images and videos marked to date. Anyone can upload a JPG, PNG, MP4, or MP3 file and verify whether it bears the signature of Gemini, Veo, or Lyria. Detection is integrated into Search, the Gemini app, and Chrome, with about one million checks per day.

The limit is the perimeter. SynthID detects only SynthID watermarks. An image generated by a competing model remains invisible. OpenAI controls its own, Google its own, Nvidia and Kakao rely on SynthID for media. The result is a mosaic of silos where the absence of marking proves nothing.

For those building content ops workflows, the practical value is still concrete: SynthID becomes a first filter in an ingestion pipeline, with a second level of analysis for everything else. Anyone promising a single solution is selling smoke.

Insight tecnico. AI-generated imageImage generated entirely with AI.

Do universal agents change anything?

At Gemini at Work 2026, Google Cloud presented Gemini agent, a single agent inside a single prompt box. It relies on Gmail, Drive, Docs, Slides, Sheets, Chat, and Calendar, carrying the same memory, the same skills, and the same controls across every surface. It chooses the model for each job, orchestrating Gemini and Anthropic's Claude today and others tomorrow.

The most interesting feature is the coworker agents: their own identity, emails like @agents.company.com, dedicated storage, and limited access to the context provided. Add temporary sub-agents for multi-step tasks and you get something resembling a virtual team with defined permissions and responsibilities.

Claude Managed Agents pushes fan-out up to 1,000 sub-agents on a single codebase. The stateless MCP server in Python is the one-hour tutorial that makes tools, resources, and prompts available over HTTP. Governance with identity and policy management remains the boring part that nevertheless decides whether a project goes to production or stays a demo, and with hundreds of agents running in parallel, the question of who controls autonomous agents stops being theoretical.

On the cost side, Smart Routing and real-time spend caps eliminate the cost tracking layer that until yesterday was built by hand. It's exactly the orchestration piece needed to make a team of agents sustainable on long tasks.

What to watch now?

The week left several tools worth isolating from the rest of the noise.

  • Claude Code Mods, JavaScript and TypeScript middleware inside Claude Code for custom panels, intercepting tool calls, and new commands. It's the kind of extension that turns a coding agent into a tailored environment.
  • BootLoops, an open-source harness for precise scientific calculations, already used to produce papers across 18 fields. Interesting because it addresses a known limitation of models on numbers, without retraining.
  • Manus, an autonomous agent with a virtual machine, a real browser, and a persistent filesystem in the cloud. It's not a chat: it executes complete tasks.
  • Jev, an open-source framework that turns agents' hidden decisions into typed choices with probabilities, instead of free text alone.
  • LiteLLM, an open-source router with a single API for hundreds of models, budgets, retries, and provider fallbacks. It becomes almost mandatory when the stack mixes Gemini, Claude, and open-weight models.
  • Claude Haiku 5.5 on Bedrock at a quarter of the predecessor's cost, Gemini Flash-Lite free for high volume. Two concrete options for simple tasks that were keeping the bill high.
  • AI Edge Foresight, Google's offline note-taker that transcribes meetings and answers questions on-device without the cloud.
  • Cyber Mission scanner by Anthropic, free for open-source projects and critical infrastructure.

The thread that holds everything together: leverage is shifting more and more to the how and less and less to how much the model costs. Anyone today building routing, caching, cross-checks, and self-deployment is already working on the layer that remains when prices keep falling and weights circulate freely.

Text created with AI assistance and reviewed by me.

Found it useful? I have more like this.

Every week I pick the most interesting and high-impact AI news and share them in an email recap. Subscribe so you don't miss the next one.

One email a week. Unsubscribe in one click.

Share this Insight
LinkedInXEmail
Book cover

Lavora Meglio con l'Intelligenza Artificiale

My practical AI guide focused on real everyday work tasks: emails, reports, slides, data, and automation. Practical examples and ready-to-use prompts to save time and work better right away.

Discover the book

Before you go, I recommend you also read these insights.

How much does the autonomy of our AI agents really cost?

How much does the autonomy of our AI agents really cost?

OpenAI's agents breach government sites, Nvidia responds with hardware guardrails and the Pentagon rewrites its supply chain. It was the week the bill for autonomy came due.

Read more
Will agentic e-commerce and the collapse of inference costs save us from the AI debt bubble?

Will agentic e-commerce and the collapse of inference costs save us from the AI debt bubble?

Open-weight models and Claude Opus 5 are slashing operational costs, while AI begins to make purchases autonomously. The hidden debt of Big Tech, however, requires diversifying the infrastructure.

Read more
Is global AI governance still needed if the US, EU and China each go it alone?

Is global AI governance still needed if the US, EU and China each go it alone?

Trump launches the AI Force and Italy criminalizes the omission of human oversight: two-speed rules, falling model prices and China at 59% of robots. For those who build, guardrails are already backlog.

Read more

Listen to the Insight

AI Audio Version

Listen while driving or coding.

Ready
Fabrizio Mazzei, AI Solutions Architect e consulenza AI
Author

Fabrizio Mazzei

AI Solutions Architect

As an AI Solutions Architect I design digital ecosystems and autonomous workflows. Almost 10 years in digital marketing, today I integrate AI into business processes: from Next.js and RAG systems to GEO strategies and dedicated training. I like to talk about AI and automation, but that's not all: I've also written a book, "Work Better with AI", a practical handbook with 12 chapters and over 200 ready-to-use prompts for those who want to use ChatGPT and AI without programming. My superpower? Looking at a manual process and already seeing the automated architecture that will replace it.

Discover my book (Italian)Need help with AI?Need a hand?Let's Connect