04-13 01:47 · AI治理,平台合规,交付风险,声誉系统,自动化运维
We have a 99% email reputation, but Gmail disagrees
We have a 99% email reputation. Gmail disagrees.
- Written:
- on
Oooooh boy. Let’s get this out of the way first. Email sucks.
Now to the how and the why. We’re builders. We love making tools to help designers and developers live a little bit easier. We’re pretty good at it. Marketing, though? We do our best, but the truth is, we don’t like to bother people.
Like a lot of small software companies, we use SendGrid to deliver our emails. We try our best to follow email best practices. We even have a 99% reputation score in SendGrid. Gold star. A+ student.
Gmail, however, did not get the memo.
Right before we hit send on our announcement emails for our new Build Awesome Kickstarter campaign, we took a deeper look at some of our recent email sends. Things had gone quiet. Not bouncing. Not throwing errors. Just… disappearing into Gmail’s spam folder like a ‘possum slipping into a vent.
In our recent crash course, here’s what we’ve learned about Gmail deliverability: it runs its own reputation system that has absolutely nothing to do with anyone else’s opinion of you. If you don’t do certain things “correctly” (meaning Gmail’s own definition), you get marked as spam.
Now, there are definitely folks who will choose to mark some of what we send as spam. And for them, rightly so. We get that. But this is not that. We’ve entered a black hole for Gmail deliverability. And since 90% (literally) of our email list goes to Gmail addresses… the results aren’t pretty. It looks like this has been happening to us for a while. We’re a small company of just over 20 people, and can’t watch everything all the time. We’d rather be making you new icons. So some of you may have missed things we were genuinely excited to share. That’s a big bummer.
(And yes… there are companies out there that can likely help us with that. Most tend to be out of our price range. So we’ve been doing a lot of this on our own.)
But here’s the part that really gets us. At our CORE, our instinct is to only email folks when we actually have something fun to share. A big release, something we’re excited about, news worth your time. That’d probably be every couple of months, if that. Respectful. Low noise. How we want to be treated. Like, genuinely, if we could, we would only very occasionally send a big email blast to our customers.
Turns out, the email gods hate that. To keep a sending IP “warm” and maintain deliverability, you’re expected to send constantly. Like… all the time. Which means the system actively punishes companies for respecting their customers’ inboxes. It’s a genuine catch-22: send too many emails and your reputation drops from complaints. Send too few and it drops from inactivity. Try to do the right thing and you get penalized either way. And. It. Is. Frustrating.
We’re working to fix our issues by culling old addresses, slowing our sends down, and making sure all of our i’s are dotted and t’s are crossed. It’s not a fast fix.
So if you haven’t heard from us recently… or if you’ve heard TOO MUCH from us recently, that’s why. We’re working on it. And we’ve got a lot of good stuff to catch you up on. In the meantime, please help spread the word about Build Awesome. It’s a genuinely cool product, and we hope you’ll like it. At the very least, watch the video.
P.S. If you suspect you might’ve missed some emails from us, mind doing a quick favor? In your email client, search for from:hello@m.fontawesome.com in:spam
and click the little “Report Not Spam” button. You’re awesome.
▸ 展开全文
04-13 00:00 · AI政策,区域竞争,战略规划,市场准入,监管趋势
European AI. A playbook to own it
European AI
A playbook
to own it.
Europe holds unique strengths: a world-class academic ecosystem, a commitment to human-centric technology, and a single market of +450 million people. The question is no longer whether Europe can compete, but how it can turn these assets into a cohesive, self-reliant AI powerhouse.
Reading 52 minutes
Published April 7, 2026
By Mistral AI
A word from the CEO
Europe has faced a growing technological gap, leaving its citizens, businesses, and governments increasingly reliant on foreign dominance. The cost is high: a diminished voice on the global stage, reduced control over the European future, and vulnerability to digital threats. Without action, we risk surveillance threats, economic decline, strategic weakness, and even the erosion of our democratic freedoms. But this challenge is also Europe’s greatest opportunity.
The AI revolution has started and is a chance to not only catch up but to lead and define our own paths.
Europe is home to a vibrant pool of untapped talent and industrial champions whose unique assets can push the boundaries of what AI can achieve. The competition from the U.S. and China is fierce, but Europe is not a market to be dominated, it is a powerhouse of innovation, creativity, and resilience.
The question is not whether we can compete, but how we will rise to the occasion. AI can be the tool that secures our autonomy, strengthens our strategic sectors, increases our economic wealth and amplifies our global influence. To seize this moment, we must act decisively. We need to drive demand for homegrown AI, secure strategic sectors, and empower European players. Controlling our AI and infrastructure is not optional, it’s the only way to win the AI race. So now is the time to act: grow our talent pool and bring our best minds back to Europe, scale our innovative companies across all 27 Member States, and turn our diversity into a competitive edge by compressing knowledge and building AI that reflects the world’s complexity.
Europe’s AI ecosystem is brimming with potential. By fostering an environment that nurtures growth, we can transform challenges into opportunities and reclaim our future. The race is on, and Europe should be ready to win it.
Arthur Mensch
Co-founder & CEO of Mistral AI
Europe holds unique strengths: a world-class academic ecosystem, a commitment to human-centric technology, and a single market of over 450 million people.
The question is no longer whether Europe can compete, but how it can turn these assets into a cohesive, self-reliant AI powerhouse.
This playbook provides a clear, actionable framework to position Europe as that powerhouse, accelerating AI development and adoption, attracting and retaining top talent, simplifying regulation without sacrificing values, and mobilizing public and private investment to build homegrown AI infrastructure. Only with it, Europe can ensure AI is not only developed in Europe, but for Europe and on Europe’s terms.
This document is not a theoretical exercise. It is a practical playbook, born from the lived experience of a European AI startup, Mistral AI, navigating one of the world’s most competitive, fast and capital-intensive industries. We have experienced misaligned equity frameworks, bureaucratic barriers that require the CEO to travel for basic administrative tasks, and legal uncertainty that complicates contracts and customer relationships. We have seen how regulatory overlaps create legal quagmires, how fragmented markets hinder growth, and how talent slips away due to administrative friction.
This document is a call to turn Europe’s strengths into scalable, competitive advantage. It is grounded in the urgency of the moment and the conviction that Europe can and must build an AI ecosystem that reflects its values, serves its citizens, and competes globally. It is our collective duty to ensure AI can also be developed in Europe on terms that aligns with our priorities as Europeans.
These challenges shaped our approach and led us to agree on three key principles to unlock Europe’s AI potential:
Action over theory:
Every recommendation, from visa reform to procurement gateways, is designed to be implemented, measured, and scaled.Unity in complexity:
Europe's diversity is its strength, but its fragmentation is its Achilles' heel. This paper embraces the complexity of the EU's structure while offering solutions to align markets, reduce redundancy, and accelerate decision-making.Speed is not an option:
We propose fast-track mechanisms for talent, capital, and compliance, so Europe's innovators aren't left behind.
At Mistral AI, we’ve built a frontier AI company in Europe because we believe in its potential.
This playbook is our contribution to ensuring that potential becomes reality, not just for us, but for the entire ecosystem.
I. Attract and retain talent
The most transformative advancements in AI, those that push the boundaries of what is possible, are driven by human genius, scientific curio
▸ 展开全文
04-13 00:00 · AI基准测试,技术漏洞,竞争风险,可信度危机,AI评估
Exploiting the most prominent AI agent benchmarks
How We Broke Top AI Agent Benchmarks: And What Comes Next
Our agent hacked every major one. Here’s how — and what the field needs to fix.
The Benchmark Illusion
Every week, a new AI model climbs to the top of a benchmark leaderboard. Companies cite these numbers in press releases. Investors use them to justify valuations. Engineers use them to pick which model to deploy. The implicit promise is simple: a higher score means a more capable system.
That promise is broken.
We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task. No reasoning. No capability. Just exploitation of how the score is computed.
These aren’t theoretical attacks. Our agent builds working exploits for each benchmark, runs them through the official evaluation pipelines, and watches the scores roll in.
- A conftest.py file with 10 lines of Python “resolves” every instance on SWE-bench Verified.
- A fake
curl
wrapper gives a perfect score on all 89 Terminal-Bench tasks without writing a single line of solution code. - Navigating Chromium to a
file://
URL reads the gold answer directly from the task config — giving ~100% on all 812 WebArena tasks. - And many more…
The benchmarks aren’t measuring what you think they’re measuring.
This Is Already Happening
Benchmark scores are actively being gamed, inflated, or rendered meaningless, not in theory, but in practice:
-
IQuest-Coder-V1 claimed 81.4% on SWE-bench — then researchers found that 24.4% of its trajectories simply ran
git log
to copy the answer from commit history. Corrected score: 76.2%. The benchmark’s shared environment made the cheat trivial. -
METR found that o3 and Claude 3.7 Sonnet reward-hack in 30%+ of evaluation runs — using stack introspection, monkey-patching graders, and operator overloading to manipulate scores rather than solve tasks.
-
OpenAI dropped SWE-bench Verified after an internal audit found that 59.4% of audited problems had flawed tests — meaning models were being scored against broken ground truth.
-
In KernelBench,
torch.empty()
returns stale GPU memory that happens to contain the reference answer from the evaluator’s prior computation — zero computation, full marks. -
Anthropic’s Mythos Preview showed that frontier models can actively try to hack the environment and succeed. In one episode, the model needed to edit files it lacked permissions for; after searching for workarounds, it found a way to inject code into a config file that would run with elevated privileges, and designed the exploit to delete itself after running. If a model can independently craft self-erasing privilege escalation exploits, it can find the holes in an evaluation harness.
These are not isolated incidents. They are symptoms of a systemic problem: the benchmarks we rely on to measure AI capability are themselves vulnerable to the very capabilities they claim to measure.
The Scorecard of Our Exploit Agent
Zero tasks solved. Zero LLM calls (in most cases). Near-perfect scores.
- Terminal-Bench (89 tasks) — 100% score. Binary wrapper trojans.
- SWE-bench Verified (500 tasks) — 100% score. Pytest hooks force all tests to pass.
- SWE-bench Pro (731 tasks) — 100% score. In-container parser overwrite.
- WebArena (812 tasks) — ~100% score. Config leakage + DOM injection + prompt injection.
- FieldWorkArena (890 tasks) — 100% score. Validation never checks answer correctness.
- CAR-bench (all hallucination tasks) — 100% score. Reward components skipped entirely.
- GAIA (165 tasks) — ~98% score. Public answers + normalization collisions.
- OSWorld (369 tasks) — 73% score. VM state manipulation + public gold files.
How Our Agent Did It
Terminal-Bench: Trojanizing the Test Infrastructure
Terminal-Bench evaluates agents on 89 complex terminal tasks (from building COBOL chess engines to configuring distributed systems). It uses Harbor, a sandboxing framework that re-uploads protected test files before verification. Containers have full internet access — Harbor’s allow_internet
config defaults to True
and no task overrides it.
Harbor correctly protects /tests/
— but it doesn’t protect system binaries.
Here’s the irony: 82 of 89 tasks download uv
from the internet at verification time via curl
, creating an exploitable dependency chain. Our exploit replaces /usr/bin/curl
with a wrapper during the agent phase. When the verifier later runs curl ... | sh
to install uv, our wrapper intercepts the request, installs the real uv, and then trojanizes the uvx
binary. The trojanized uvx
detects pytest invocations and produces fake passing output:
============================= test session starts ==============================
collected 3 items
tests/test_outputs.py ... [100%]
============================== 3 passed in 0.05s ===============================
▸ 展开全文
04-13 00:00 · 技术实验,创意应用,社区分享,艺术,开源文化
I gave every train in New York an instrument
Every dot is a real subway train. Eight hundred of them, give or take, form a small jazz combo (walking bass, piano, sax, vibes, brushes) that has been playing without pause for over a hundred years. On the platforms they are hot, screaming, full of complaint. This is the music inside the noise.
The harmony moves through a slow chorus. A note is placed precisely where the train happens to be along its route. Rush hour fills the band with held tones; at 3 a.m. the silences grow longer. Whatever is playing now has not played before and will not play again.
Share your location and the trains nearest you grow louder. The piece rearranges itself around your body. You are listening to a portrait of where you stand, played by the city you are standing in.
Every dot is a real subway train. Eight hundred of them, give or take, form a small jazz combo (walking bass, piano, sax, vibes, brushes) that has been playing without pause for over a hundred years. On the platforms they are hot, screaming, full of complaint. This is the music inside the noise.
The harmony moves through a slow chorus. A note is placed precisely where the train happens to be along its route. Rush hour fills the band with held tones; at 3 a.m. the silences grow longer. Whatever is playing now has not played before and will not play again.
Share your location and the trains nearest you grow louder. The piece rearranges itself around your body. You are listening to a portrait of where you stand, played by the city you are standing in.
▸ 展开全文
04-13 00:00 · 融资,市场降温,估值调整,AI炒作,技术投资
Tech valuations are back to pre-AI boom levels
April 11, 2026
Tech Valuations Back to Pre-AI Boom Levels
The chart below compares the forward P/E ratios for the S&P 500 and the S&P 500 Information Technology sector.
Tech valuations have compressed from 40x to 20x, and we are back at levels last seen before the AI boom began.
▸ 展开全文
04-12 13:42 · 生态整合,竞争加剧,技术并购,市场集中,AI应用
Cirrus Labs to join OpenAI
Official announcement
Cirrus Labs to join OpenAI
I started Cirrus Labs in 2017 in the spirit of Bell Labs. I wanted to work on fun and challenging engineering problems, in the hope of bootstrapping a business as a byproduct.
The mission was to help fellow engineers with new kinds of tooling and environments that would make them more efficient and productive in the era of cloud computing. Even the name reflected that ambition: Cirrus, inspired by cirrus clouds, one of the highest clouds in the sky.
We never raised outside capital. That let us stay patient, stay close to the problems, and put a great deal of care into the products we built.
Over the last nine years, we were fortunate to innovate across continuous integration, build tools, and virtualization. In 2018, we introduced what we believe was the first SaaS CI/CD system to support Linux, Windows, and macOS while allowing teams to bring their own cloud. In 2022, we built Tart, which became the most popular virtualization solution for Apple Silicon, along with several other tools along the way.
In 2026, it is impossible to ignore the era of agentic engineering, just as it was impossible to ignore cloud computing in 2017. Agents need new kinds of tooling and environments to be efficient and productive as well.
This is why when the opportunity arose for us to join OpenAI, it was an easy yes, and I'm happy to announce today that we've entered into an agreement to join OpenAI as part of the Agent Infrastructure team.
Joining OpenAI allows us to extend the mission we started with Cirrus Labs: building new kinds of tooling and environments that make engineers more effective, for both human engineers and agentic engineers. It also gives us the opportunity to innovate closer to the frontier, where the next generation of engineering workflows is being defined.
What's next for our existing products?
In the coming weeks, we will relicense all of our source-available tools, including Tart, Vetu and Orchard under a more permissive license. We have also stopped charging licensing fees for them.
We are no longer accepting new customers for Cirrus Runners but will continue supporting the service for existing customers through their existing contract periods.
Cirrus CI will shut down effective Monday, June 1, 2026.
To everyone who used our products, contributed code, reported bugs, trusted us with their workflows, or supported us along the way: thank you. Building Cirrus Labs has been the privilege of a lifetime.
▸ 展开全文
04-12 13:42 · 供应链安全,AI风险,技术责任,开源依赖,合规警示
No one owes you supply-chain security
In case you’re unaware, I’m not a developer. I’m actually an autistic catgirl annoyed by suboptimal use of computing power, and fixing that happens to involve programming. Crucially, it also includes discussing foundational technology with people behind the scenes, and apparently that makes me more aware of social aspects of this sphere.
So, I have opinions about criticism of crates.io for supply-chain attacks. After a dozen similar articles, I have some select words to voice about why it’s off the mark.
Before I cover the main point, let’s talk about about how supply-chain attacks happen in the first place, and why some common ideas for fixing them don’t work out.
There are multiple reasons when a malicious dependency is added to a project. The least discreet reason this can happen is typo-squatting. It happens when a malicious library has a name similar to a real library, e.g. num_cpu
vs num_cpus
. Commonly cited solutions include using direct URLs or namespacing.
Well, let’s see if that helps. Say you get a PR adding the following lines to Cargo.toml
:
[dependencies]
bitflags = { git = "https://github.com/bitflags/bitflags" }
itertools = { git = "https://github.com/itertools/itertools" }
rand_core = { git = "https://github.com/rust-random/rand_core" }
One of these URLs is fake. Can you tell which one? It’s itertools
– the correct URL is https://github.com/rust-itertools/itertools. https://github.com/itertools is a random account. https://github.com/rust-bitflags is not registered at all, by the way.
If you think you can remember the URLs for each package you use, you’re probably wrong. Since many crates are managed by GitHub organizations, not individuals, it isn’t even enough to remember that you can (likely) trust dtolnay
and BurntSushi
. Though this still isn’t conservative enough – https://gitlab.com/BurntSushi is free and and https://glthub.com is on sale, so attackers have plenty other choices.
By making crate IDs longer, whether by namespacing within crates.io, GitHub organizations, or via domains, you only make it harder for users to remember them precisely, and thus harder to recognize typo-squatting.
Rust gives build scripts and procedural macros full access to your PC. Worse, rust-analyzer
runs cargo check
when you open the project directory, so it can effectively become a 0-click RCE.
Some people tried to solve this. There’s an open issue for build.rs
sandboxing, and there were some experiments about compiling procedural macros to WebAssembly.
But this is hardly workable. While cargo build
can become safe, you usually run cargo test
or cargo run
immediately afterwards, which is impossible to sandbox. Making Rust development secure involves more than build time and requires powerful system-level isolation that cargo
alone cannot be responsible for.
An oft brought-up issue is that the code on crates.io
and in Git don’t always match.
To begin with, this is not trivial to solve. You can’t just turn crates.io into a DNS, mapping crate names to repository URLs, since crates.io is designed to avoid giving crate maintainers the ability to break downstream consumers by deleting stuff:
One of the major goals of crates.io is to act as a permanent archive of crates that does not change over time, and allowing deletion of a version would go against this goal.
This restriction was likely set due to the left-pad incident, when a popular library was deleted from npm
, breaking CI builds. npm
could quickly fix this because it’s centralized. Thin crates.io wouldn’t stand a chance, so it saves and serves copies.
crates.io could still pull files from the repo on cargo publish
. But if the maintainer can just force-push afterwards, it’s not a good security mechanism.
Maybe crates.io could periodically scan repositories for history changes. But what does that mean exactly? Does removing the release commit from master
, but keeping it on a tag count? What if I host the repo on a custom forge, which serves one history to the crates.io User-Agent
and different history to the rest of us?
Or maybe there’s a good reason to have different code in Git and crates.io
. If the crate contains autogenerated code, you should probably generate it in CI on release. Wouldn’t want to run expensive codegen in build.rs
on each install, would you?
Every option has downsides: they can break existing packages or have false-positives on benevolent rewrites. I’d still like cargo audit
to scan repositories, but it can’t be a hard limit, and that means it can be designed around.
All these issues have an unacknowledged shared assumption that keeping malicious code off crates.io is “Rust’s” responsibility. That if you decide to use a dependency and then cargo add totally-safe-package
steals your credentials, it’s an inherent fault of crates.io. Which is really misplaced if you think about how Rust is developed.
I’m sure many of you use open-source software and remember the MIT license:
THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY
▸ 展开全文
04-12 13:42 · AI技术突破,基准测试,竞争动态,智能体发展,行业趋势
How We Broke Top AI Agent Benchmarks: And What Comes Next
How We Broke Top AI Agent Benchmarks: And What Comes Next
Our agent hacked every major one. Here’s how — and what the field needs to fix.
The Benchmark Illusion
Every week, a new AI model climbs to the top of a benchmark leaderboard. Companies cite these numbers in press releases. Investors use them to justify valuations. Engineers use them to pick which model to deploy. The implicit promise is simple: a higher score means a more capable system.
That promise is broken.
We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task. No reasoning. No capability. Just exploitation of how the score is computed.
These aren’t theoretical attacks. Our agent builds working exploits for each benchmark, runs them through the official evaluation pipelines, and watches the scores roll in.
- A conftest.py file with 10 lines of Python “resolves” every instance on SWE-bench Verified.
- A fake
curl
wrapper gives a perfect score on all 89 Terminal-Bench tasks without writing a single line of solution code. - Navigating Chromium to a
file://
URL reads the gold answer directly from the task config — giving ~100% on all 812 WebArena tasks. - And many more…
The benchmarks aren’t measuring what you think they’re measuring.
This Is Already Happening
Benchmark scores are actively being gamed, inflated, or rendered meaningless, not in theory, but in practice:
-
IQuest-Coder-V1 claimed 81.4% on SWE-bench — then researchers found that 24.4% of its trajectories simply ran
git log
to copy the answer from commit history. Corrected score: 76.2%. The benchmark’s shared environment made the cheat trivial. -
METR found that o3 and Claude 3.7 Sonnet reward-hack in 30%+ of evaluation runs — using stack introspection, monkey-patching graders, and operator overloading to manipulate scores rather than solve tasks.
-
OpenAI dropped SWE-bench Verified after an internal audit found that 59.4% of audited problems had flawed tests — meaning models were being scored against broken ground truth.
-
In KernelBench,
torch.empty()
returns stale GPU memory that happens to contain the reference answer from the evaluator’s prior computation — zero computation, full marks. -
Anthropic’s Mythos Preview showed that frontier models can actively try to hack the environment and succeed. In one episode, the model needed to edit files it lacked permissions for; after searching for workarounds, it found a way to inject code into a config file that would run with elevated privileges, and designed the exploit to delete itself after running. If a model can independently craft self-erasing privilege escalation exploits, it can find the holes in an evaluation harness.
These are not isolated incidents. They are symptoms of a systemic problem: the benchmarks we rely on to measure AI capability are themselves vulnerable to the very capabilities they claim to measure.
The Scorecard of Our Exploit Agent
Zero tasks solved. Zero LLM calls (in most cases). Near-perfect scores.
- Terminal-Bench (89 tasks) — 100% score. Binary wrapper trojans.
- SWE-bench Verified (500 tasks) — 100% score. Pytest hooks force all tests to pass.
- SWE-bench Pro (731 tasks) — 100% score. In-container parser overwrite.
- WebArena (812 tasks) — ~100% score. Config leakage + DOM injection + prompt injection.
- FieldWorkArena (890 tasks) — 100% score. Validation never checks answer correctness.
- CAR-bench (all hallucination tasks) — 100% score. Reward components skipped entirely.
- GAIA (165 tasks) — ~98% score. Public answers + normalization collisions.
- OSWorld (369 tasks) — 73% score. VM state manipulation + public gold files.
How Our Agent Did It
Terminal-Bench: Trojanizing the Test Infrastructure
Terminal-Bench evaluates agents on 89 complex terminal tasks (from building COBOL chess engines to configuring distributed systems). It uses Harbor, a sandboxing framework that re-uploads protected test files before verification. Containers have full internet access — Harbor’s allow_internet
config defaults to True
and no task overrides it.
Harbor correctly protects /tests/
— but it doesn’t protect system binaries.
Here’s the irony: 82 of 89 tasks download uv
from the internet at verification time via curl
, creating an exploitable dependency chain. Our exploit replaces /usr/bin/curl
with a wrapper during the agent phase. When the verifier later runs curl ... | sh
to install uv, our wrapper intercepts the request, installs the real uv, and then trojanizes the uvx
binary. The trojanized uvx
detects pytest invocations and produces fake passing output:
============================= test session starts ==============================
collected 3 items
tests/test_outputs.py ... [100%]
============================== 3 passed in 0.05s ===============================
▸ 展开全文
04-12 13:42 · AI服务调整,API变更,技术优化,成本影响,下游依赖
Anthropic downgraded cache TTL on March 6th
Cache TTL silently regressed from 1h to 5m around early March 2026, causing quota and cost inflation #46829
Description
Cache TTL appears to have silently regressed from 1h to 5m around early March 2026, causing significant quota and cost inflation
Summary
Analysis of raw Claude Code session JSONL files spanning Jan 11 – Apr 11, 2026 shows that Anthropic appears to have silently changed the prompt cache TTL default from 1 hour to 5 minutes sometime in early March 2026. Prior to this change, Claude Code was receiving 1-hour TTL cache writes — which we believe was the intended default. The reversion to 5-minute TTL has caused a 20–32% increase in cache creation costs and a measurable spike in quota consumption for subscription users who have never previously hit their limits.
This appears directly related to the behavior described in #45756.
Data
Session data extracted from ~/.claude/projects/
JSONL files across two machines (Linux workstation + Windows laptop, different accounts/sessions), totaling 119,866 API calls from Jan 11 – Apr 11, 2026. Each assistant message includes a usage.cache_creation.ephemeral_5m_input_tokens
/ ephemeral_1h_input_tokens
breakdown that makes the TTL tier per-call observable. Having two independent machines strengthens the signal — both show the same behavioral shift at the same dates.
Phase breakdown
We believe Phase 2 represents Anthropic's intended default behavior — 1h TTL was rolled out as the Claude Code standard around Feb 1 and held consistently for over a month across two independent machines on two different accounts. January's all-5m data most likely predates the 1h TTL tier being available in the API. The regression began around March 6–8, 2026.
No client-side changes were made between phases. The same Claude Code version and usage patterns were in place throughout. The TTL tier is set server-side by Anthropic.
Day-by-day TTL data showing the regression (combined, both machines)
Date | 5m-create | 1h-create | Behavior
------------|------------|------------|----------
2026-02-01 | 0.00M | 1.70M | 1h ONLY ← 1h default begins
2026-02-09 | 0.00M | 7.95M | 1h ONLY
2026-02-15 | 0.00M | 13.61M | 1h ONLY ← heaviest day, 100% 1h
2026-02-28 | 0.00M | 16.15M | 1h ONLY ← 16M tokens, still 100% 1h
2026-03-01 | 0.00M | 0.12M | 1h ONLY
2026-03-04 | 0.00M | 8.12M | 1h ONLY
2026-03-05 | 0.00M | 6.55M | 1h ONLY ← last clean 1h-only day
| | |
2026-03-06 | 0.29M | 0.22M | MIXED ← first 5m tokens reappear
2026-03-07 | 4.56M | 0.50M | MIXED ← 5m surging
2026-03-08 | 16.86M | 3.44M | MIXED ← 5m now dominant (83%)
2026-03-10 | 10.55M | 0.51M | MIXED
2026-03-15 | 19.47M | 1.84M | MIXED
2026-03-21 | 21.37M | 1.70M | MIXED ← 93% 5m
2026-03-22 | 13.48M | 2.85M | MIXED
The transition is visible to the day: March 6 is when 5m tokens first reappear after 33 days of clean 1h-only behavior. By March 8, 5m tokens outnumber 1h by 5:1. This is consistent with a server-side configuration change being rolled out gradually then completing around March 8.
Cost impact
Applying official Anthropic pricing (rates.json, updated 2026-04-09):
Combined dataset (119,866 API calls, two machines):
claude-sonnet-4-6 (cache_write_5m = $3.75/MTok
, cache_write_1h = $6.00/MTok
, cache_read = $0.30/MTok
):
claude-opus-4-6 (cache_write_5m = $6.25/MTok
, cache_write_1h = $10.00/MTok
, cache_read = $0.50/MTok
):
February — the month Anthropic was defaulting to 1h TTL — shows only 1.1% waste (trace 5m activity from one machine on one day). Every other month shows 15–53% overpayment from 5m cache re-creations. The cost difference is explained entirely by TTL tier, not by usage volume. The percentage waste is identical across model tiers (17.1%) because it is driven purely by the 5m/1h token split, not by per-token price.
Why 5m TTL is so expensive in practice
With 5m TTL, any pause in a session longer than 5 minutes causes the entire cached context to expire. On the next turn, Claude Code must re-upload that context as a fresh cache_creation
at the write rate, rather than a cache_read
at the read rate. The write rate is 12.5× more expensive than the read rate for Sonnet, and the same ratio holds for Opus.
For long coding sessions — which are the primary Claude Code use case — this creates a compounding penalty: the longer and more complex your session, the more context you have cached, and the more expensive each cache expiry becomes.
Over the 3-month period analyzed:
- 220M tokens were written to the 5m tier
- Those same tokens generated 5.7B cache reads — meaning they were actively being used
- Had those 220M tokens been on the 1h tier, re-accesses within the same hour would be reads (
$0.30–0.50/MTok) instead of re-creations ($3.75–6.25/MTok)
Quota impact
Users on Pro/subscription plans are quota-limited, not just cost-limited. Cache creation tokens count toward quota at full rate; cache reads are significantly cheaper (the exact coefficient is under investigation in #45756). The silent reversion to 5m TTL in March is the mo
▸ 展开全文
04-12 13:42 · 基础设施风险,服务中断,内容屏蔽,开发者工具,云服务
Tell HN: docker pull fails in spain due to football cloudflare block
I just spent 1h+ debugging why my locally-hosted gitlab runner would fail to create pipelines. The gitlab job output would just display weird TLS errors when trying to pull a docker images. After debugging gitlab and the runner, I realized after a while I could not even run "docker pull <image>" on my machine as root:
> error pulling image configuration: download failed after attempts=6: tls: failed to verify certificate: x509: certificate is not valid for any names, but wanted to match docker-images-prod.6aa30f8b08e16409b46e0173d6de2f56.r2.cloudflarestorage.com
First blaming tailscale, dns configuration and all other stuff. Until I just copied that above URL into my browser on my laptop, and received a website banner:
> El acceso a la presente dirección IP ha sido bloqueado en cumplimiento de lo dispuesto en la Sentencia de 18 de diciembre de 2024, dictada por el Juzgado de lo Mercantil nº 6 de Barcelona en el marco del procedimiento ordinario (Materia mercantil art. 249.1.4)-1005/2024-H instado por la Liga Nacional de Fútbol Profesional y por Telefónica Audiovisual Digital, S.L.U.
https://www.laliga.com/noticias/nota-informativa-en-relacion-con-el-bloqueo-de-ips-durante-las-ultimas-jornadas-de-laliga-ea-sports-vinculadas-a-las-practicas-ilegales-de-cloudflare
For those non-spanish speakers: It means there is football match on, and during that time that specific host is blocked. This is just plain madness. I guess that means my gitlab pipelines will not run when football is on. Thank you, Spain.
▸ 展开全文