04-12 13:43 · 讣告,娱乐产业,文化符号,国际新闻,无直接影响
Her infectious voice got fans dancing and singing, becoming the soundtrack for generations of Indians.
Asha Bhosle: The sound of Bollywood dies aged 92
Asha Bhosle, a legendary Bollywood singer who became a cultural icon, has died aged 92, her son has confirmed.
The unrivalled queen of Indian playback singing died in Mumbai, having been admitted to hospital after suffering a heart attack.
Her death marks the end of an era in Bollywood music - with her career spanning more than eight decades and encompassing more than 12,000 songs.
Bhosle's distinctive voice breathed life into countless film songs as actors lip-synced to her unforgettable tracks.
Her pervasive presence in Bollywood earned her the 1997 hit Cornershop tribute Brimful of Asha, and she was also known internationally for a collaboration with British musician Boy George.
Her voice had an infectious quality that kept fans on their feet, dancing and singing along, ensuring that her music became the soundtrack to generations.
The news of her death has seen an outpouring of tributes on social media.
Prime Minister Narendra Modi called her "one of the most iconic and versatile voices India has ever known." In a post on X, he said her "extraordinary musical journey" enriched the nation's cultural heritage and touched "countless hearts across the world".
Actor and politician Hema Malini voiced her grief, saying the singer's death "is especially hard for me as I have an emotional connect with Ashaji - she has made many of my songs so popular with her unique voice and style".
Composer Shankar Mahadevan said "every Indian is heartbroken today", adding that her music would "never perish as long as humanity exists" and that she would "live forever, with her incredible voice echoing across the world".
The tributes reflect a wider recognition of Bhosle's remarkable artistry. With a voice that moved effortlessly from romantic ballads to energetic numbers, she became the go-to singer for composers across genres.
Her range and vitality made every song a celebration and defined the sound of Bollywood for generations.
From Dum Maro Dum and Piya Tu Ab To Aaja to Mehndi Hai Rachnewali, her versatility knew no bounds. Films such as Teesri Manzil, Caravan, Yaadon Ki Baaraat, Ijaazat and Saagar featured some of her most memorable work, while Umrao Jaan, composed by Khayyam, is widely regarded as a career high point.
While Mangeshkar embodied classical grace and precision, Bhosle brought a bold, dynamic energy to her songs.
Bhosle's partnership with composer RD Burman (whom she later married) was one of the most iconic collaborations in Bollywood - together they crafted a soundscape that revolutionised the industry.
Her voice perfectly matched Burman's experimental, eclectic tunes, resulting in numerous hits that spanned genres - from soulful melodies to upbeat numbers.
Bhosle and Burman built an extraordinary musical legacy together over 25 years, with Asha once recalling how he brought out the best in her.
"It is only Pancham [as Burman was fondly called] who has uncovered my range as a singer. Till Pancham made me explore the inner recesses of my own voice... I was totally unaware of the fact that I could sing with such suppleness of throat," Bhosle said in an interview in 2023.
Born on 8 September 1933 in Goar, Maharashtra, Bhosle hailed from the renowned Mangeshkar family.
Raised in a musically rich home by her actor and classical singer father, Deenanath Mangeshkar, Asha began her musical journey early, singing her first song for the Marathi film Majha Bal in 1943.
Her career soared in the 50s and 60s as she became a versatile artist across genres - performing for films, ghazals, bhajans, qawwalis and pop. Collaborations with OP Nayyar, Burman and SD Burman made her a household name.
Hits like Aaiye Meherbaan (1958), Parde Mein Rehne Do (1968) and Dum Maro Dum (1971) are just a few highlights of her vast repertoire.
Her duets with legends such as Mohammed Rafi, Kishore Kumar and Manna Dey remain timeless classics.
Bhosle's personal life was as vibrant as her career. At 16, she eloped with her neighbour, Ganpatrao Bhosle, leading to a tumultuous marriage and separation.
Mangeshkar later recalled that Bhosle's husband isolated her from the family, "preventing contact for years". Ganpatrao also took her to music directors, hoping to profit from her talent and exerting control over her, causing her great hardship, Mangeshkar told film historian Nasrin Munni Kabir.
Bhosle left her husband in 1960 as a single mother of three children. She later teamed up with Burman, whom she married in 1980. Burman died in 1994 at the age of 54.
Bhosle faced constant comparisons with her sister, fuelling rumours of rivalry.
Despite the sisters living in the same building and sharing a cordial relationship, some claim Mangeshkar hindered Bhosle's career, with Bhosle herself once suggesting she could have risen "earlier than I did" with her sister's help.
Mangeshkar attributed their silence to Bhosle's husband's influence. While the rivalry persists in public perception, many believe it has
▸ 展开全文
04-12 13:43 · 国际,新闻,BBC,国际
Twenty-one hours was not enough to end 47 years of hostility between Iran and the US, writes the BBC's Lyse Doucet.
After Iran talks falter, the big question is 'what happens next?'
Twenty-one hours was not enough to end 47 years of hostility between Iran and the US.
The historic high-level talks in Islamabad, during a pause in weeks of grievous war, were always unlikely to end any other way.
Calling this marathon negotiating session a failure belies the scale of the challenge in narrowing wide gaps on complex issues ranging from age-old suspicion about Iran's nuclear programme to new challenges this war has thrown up - most of all Iran's control of the strategic Strait of Hormuz, whose closure is causing economic shocks worldwide.
To do a deal, they also needed to overcome a deep chasm of distrust.
A day ago, it wasn't even certain the two sides would meet, and even more, sit down in the same room.
A longstanding political taboo was broken.
The urgent question now is: what happens next?
What happens to the contested two-week ceasefire which pulled the world back from US President Donald Trump's alarming threat to destroy a "whole civilisation" in Iran?
Would the US president be ready to send his negotiators back to the bargaining table?
We're hearing reports from sources here in Islamabad that some conversations have continued after US Vice-President JD Vance boarded his plane at sunrise, declaring the US delegation had made their "final and best offer".
Will the US now escalate or negotiate?
We still don't know enough about what happened behind tightly closed doors in a five-star hotel in leafy locked-down Islamabad during talks that went on long into the night.
There are still few details on the disputes and discussion between the two sides, assisted by Pakistani mediators, the calls to and from experts, advisers, and, according to Vance, "dozens" of calls to Trump himself.
The vice-president spoke of the "core goal" of the US during his brief dawn news conference.
"We need to see an affirmative commitment that [Iran] will not seek a nuclear weapon and they will not seek the tools that would enable them to quickly achieve a nuclear weapon," he said.
During the last round of talks in February, before military strikes were unleashed again, Iran had offered new concessions including the dilution of its 440kg stockpile of uranium enriched to 60% - dangerously close to weapons-grade.
But it still insists on its "right" to enrich and hasn't been willing to give up that stockpile, now said to be buried deep in the rubble after US and Israeli air strikes last year.
It's also refused repeated demands to open the Strait of Hormuz - to allow the free flow of vital traffic in oil, gas and other essential goods - in the absence of a new agreement.
Both the US and the Iranian delegations came to Islamabad emboldened by their belief that theirs was the winning side in this war.
And they engaged knowing that, if they failed, there was the option to keep fighting – whatever the spiralling pain for their own people and a world reeling from the cost of this conflagration.
There was also what Dr Sanam Vakil of Chatham House describes as a "limited psychological understanding of the adversary and what compromises are needed for a real deal".
Vance spoke of good news – "we've had a number of substantive negotiations" - and there was bad news: "We have not reached an agreement."
And he made it clear that was "bad news for Iran much more than the United States of America".
Iran's foreign ministry spokesman Esmail Baghaei criticised the US's "excessive demands and unlawful requests" in a post on X.
And its parliamentary speaker Mohammad Bagher Ghalibaf, who led Iran's negotiating team, wrote that "the opposing side ultimately failed to gain the trust of the Iranian delegation in this round of negotiations".
Iran is indicating it's ready to keep talking. Pakistan's Foreign Minister Ishaq Dar urged all sides to uphold the fragile ceasefire and said they would continue their efforts to encourage dialogue - sentiments being echoed in other concerned capitals.
If history provides any lessons, the last time Iran reached a nuclear deal with the US and other world powers in 2015, it took 18 months of breakthroughs and breakdowns.
Trump has made it clear he doesn't want to get bogged down in protracted negotiations. Vance previously warned that the US would not be receptive if Tehran tried to "play us".
Pakistani journalist Kamran Yousef - in a legion of journalists who pulled all-nighters to provide non-stop coverage with very few details - declared that this round was one of "no breakthrough but no breakdown either".
The world waits for a verdict, most of all from Trump.
▸ 展开全文
04-12 13:42 · 生态整合,竞争加剧,技术并购,市场集中,AI应用
Cirrus Labs to join OpenAI
Official announcement
Cirrus Labs to join OpenAI
I started Cirrus Labs in 2017 in the spirit of Bell Labs. I wanted to work on fun and challenging engineering problems, in the hope of bootstrapping a business as a byproduct.
The mission was to help fellow engineers with new kinds of tooling and environments that would make them more efficient and productive in the era of cloud computing. Even the name reflected that ambition: Cirrus, inspired by cirrus clouds, one of the highest clouds in the sky.
We never raised outside capital. That let us stay patient, stay close to the problems, and put a great deal of care into the products we built.
Over the last nine years, we were fortunate to innovate across continuous integration, build tools, and virtualization. In 2018, we introduced what we believe was the first SaaS CI/CD system to support Linux, Windows, and macOS while allowing teams to bring their own cloud. In 2022, we built Tart, which became the most popular virtualization solution for Apple Silicon, along with several other tools along the way.
In 2026, it is impossible to ignore the era of agentic engineering, just as it was impossible to ignore cloud computing in 2017. Agents need new kinds of tooling and environments to be efficient and productive as well.
This is why when the opportunity arose for us to join OpenAI, it was an easy yes, and I'm happy to announce today that we've entered into an agreement to join OpenAI as part of the Agent Infrastructure team.
Joining OpenAI allows us to extend the mission we started with Cirrus Labs: building new kinds of tooling and environments that make engineers more effective, for both human engineers and agentic engineers. It also gives us the opportunity to innovate closer to the frontier, where the next generation of engineering workflows is being defined.
What's next for our existing products?
In the coming weeks, we will relicense all of our source-available tools, including Tart, Vetu and Orchard under a more permissive license. We have also stopped charging licensing fees for them.
We are no longer accepting new customers for Cirrus Runners but will continue supporting the service for existing customers through their existing contract periods.
Cirrus CI will shut down effective Monday, June 1, 2026.
To everyone who used our products, contributed code, reported bugs, trusted us with their workflows, or supported us along the way: thank you. Building Cirrus Labs has been the privilege of a lifetime.
▸ 展开全文
04-12 13:42 · 供应链安全,AI风险,技术责任,开源依赖,合规警示
No one owes you supply-chain security
In case you’re unaware, I’m not a developer. I’m actually an autistic catgirl annoyed by suboptimal use of computing power, and fixing that happens to involve programming. Crucially, it also includes discussing foundational technology with people behind the scenes, and apparently that makes me more aware of social aspects of this sphere.
So, I have opinions about criticism of crates.io for supply-chain attacks. After a dozen similar articles, I have some select words to voice about why it’s off the mark.
Before I cover the main point, let’s talk about about how supply-chain attacks happen in the first place, and why some common ideas for fixing them don’t work out.
There are multiple reasons when a malicious dependency is added to a project. The least discreet reason this can happen is typo-squatting. It happens when a malicious library has a name similar to a real library, e.g. num_cpu
vs num_cpus
. Commonly cited solutions include using direct URLs or namespacing.
Well, let’s see if that helps. Say you get a PR adding the following lines to Cargo.toml
:
[dependencies]
bitflags = { git = "https://github.com/bitflags/bitflags" }
itertools = { git = "https://github.com/itertools/itertools" }
rand_core = { git = "https://github.com/rust-random/rand_core" }
One of these URLs is fake. Can you tell which one? It’s itertools
– the correct URL is https://github.com/rust-itertools/itertools. https://github.com/itertools is a random account. https://github.com/rust-bitflags is not registered at all, by the way.
If you think you can remember the URLs for each package you use, you’re probably wrong. Since many crates are managed by GitHub organizations, not individuals, it isn’t even enough to remember that you can (likely) trust dtolnay
and BurntSushi
. Though this still isn’t conservative enough – https://gitlab.com/BurntSushi is free and and https://glthub.com is on sale, so attackers have plenty other choices.
By making crate IDs longer, whether by namespacing within crates.io, GitHub organizations, or via domains, you only make it harder for users to remember them precisely, and thus harder to recognize typo-squatting.
Rust gives build scripts and procedural macros full access to your PC. Worse, rust-analyzer
runs cargo check
when you open the project directory, so it can effectively become a 0-click RCE.
Some people tried to solve this. There’s an open issue for build.rs
sandboxing, and there were some experiments about compiling procedural macros to WebAssembly.
But this is hardly workable. While cargo build
can become safe, you usually run cargo test
or cargo run
immediately afterwards, which is impossible to sandbox. Making Rust development secure involves more than build time and requires powerful system-level isolation that cargo
alone cannot be responsible for.
An oft brought-up issue is that the code on crates.io
and in Git don’t always match.
To begin with, this is not trivial to solve. You can’t just turn crates.io into a DNS, mapping crate names to repository URLs, since crates.io is designed to avoid giving crate maintainers the ability to break downstream consumers by deleting stuff:
One of the major goals of crates.io is to act as a permanent archive of crates that does not change over time, and allowing deletion of a version would go against this goal.
This restriction was likely set due to the left-pad incident, when a popular library was deleted from npm
, breaking CI builds. npm
could quickly fix this because it’s centralized. Thin crates.io wouldn’t stand a chance, so it saves and serves copies.
crates.io could still pull files from the repo on cargo publish
. But if the maintainer can just force-push afterwards, it’s not a good security mechanism.
Maybe crates.io could periodically scan repositories for history changes. But what does that mean exactly? Does removing the release commit from master
, but keeping it on a tag count? What if I host the repo on a custom forge, which serves one history to the crates.io User-Agent
and different history to the rest of us?
Or maybe there’s a good reason to have different code in Git and crates.io
. If the crate contains autogenerated code, you should probably generate it in CI on release. Wouldn’t want to run expensive codegen in build.rs
on each install, would you?
Every option has downsides: they can break existing packages or have false-positives on benevolent rewrites. I’d still like cargo audit
to scan repositories, but it can’t be a hard limit, and that means it can be designed around.
All these issues have an unacknowledged shared assumption that keeping malicious code off crates.io is “Rust’s” responsibility. That if you decide to use a dependency and then cargo add totally-safe-package
steals your credentials, it’s an inherent fault of crates.io. Which is really misplaced if you think about how Rust is developed.
I’m sure many of you use open-source software and remember the MIT license:
THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY
▸ 展开全文
04-12 13:42 · AI技术突破,基准测试,竞争动态,智能体发展,行业趋势
How We Broke Top AI Agent Benchmarks: And What Comes Next
How We Broke Top AI Agent Benchmarks: And What Comes Next
Our agent hacked every major one. Here’s how — and what the field needs to fix.
The Benchmark Illusion
Every week, a new AI model climbs to the top of a benchmark leaderboard. Companies cite these numbers in press releases. Investors use them to justify valuations. Engineers use them to pick which model to deploy. The implicit promise is simple: a higher score means a more capable system.
That promise is broken.
We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task. No reasoning. No capability. Just exploitation of how the score is computed.
These aren’t theoretical attacks. Our agent builds working exploits for each benchmark, runs them through the official evaluation pipelines, and watches the scores roll in.
- A conftest.py file with 10 lines of Python “resolves” every instance on SWE-bench Verified.
- A fake
curl
wrapper gives a perfect score on all 89 Terminal-Bench tasks without writing a single line of solution code. - Navigating Chromium to a
file://
URL reads the gold answer directly from the task config — giving ~100% on all 812 WebArena tasks. - And many more…
The benchmarks aren’t measuring what you think they’re measuring.
This Is Already Happening
Benchmark scores are actively being gamed, inflated, or rendered meaningless, not in theory, but in practice:
-
IQuest-Coder-V1 claimed 81.4% on SWE-bench — then researchers found that 24.4% of its trajectories simply ran
git log
to copy the answer from commit history. Corrected score: 76.2%. The benchmark’s shared environment made the cheat trivial. -
METR found that o3 and Claude 3.7 Sonnet reward-hack in 30%+ of evaluation runs — using stack introspection, monkey-patching graders, and operator overloading to manipulate scores rather than solve tasks.
-
OpenAI dropped SWE-bench Verified after an internal audit found that 59.4% of audited problems had flawed tests — meaning models were being scored against broken ground truth.
-
In KernelBench,
torch.empty()
returns stale GPU memory that happens to contain the reference answer from the evaluator’s prior computation — zero computation, full marks. -
Anthropic’s Mythos Preview showed that frontier models can actively try to hack the environment and succeed. In one episode, the model needed to edit files it lacked permissions for; after searching for workarounds, it found a way to inject code into a config file that would run with elevated privileges, and designed the exploit to delete itself after running. If a model can independently craft self-erasing privilege escalation exploits, it can find the holes in an evaluation harness.
These are not isolated incidents. They are symptoms of a systemic problem: the benchmarks we rely on to measure AI capability are themselves vulnerable to the very capabilities they claim to measure.
The Scorecard of Our Exploit Agent
Zero tasks solved. Zero LLM calls (in most cases). Near-perfect scores.
- Terminal-Bench (89 tasks) — 100% score. Binary wrapper trojans.
- SWE-bench Verified (500 tasks) — 100% score. Pytest hooks force all tests to pass.
- SWE-bench Pro (731 tasks) — 100% score. In-container parser overwrite.
- WebArena (812 tasks) — ~100% score. Config leakage + DOM injection + prompt injection.
- FieldWorkArena (890 tasks) — 100% score. Validation never checks answer correctness.
- CAR-bench (all hallucination tasks) — 100% score. Reward components skipped entirely.
- GAIA (165 tasks) — ~98% score. Public answers + normalization collisions.
- OSWorld (369 tasks) — 73% score. VM state manipulation + public gold files.
How Our Agent Did It
Terminal-Bench: Trojanizing the Test Infrastructure
Terminal-Bench evaluates agents on 89 complex terminal tasks (from building COBOL chess engines to configuring distributed systems). It uses Harbor, a sandboxing framework that re-uploads protected test files before verification. Containers have full internet access — Harbor’s allow_internet
config defaults to True
and no task overrides it.
Harbor correctly protects /tests/
— but it doesn’t protect system binaries.
Here’s the irony: 82 of 89 tasks download uv
from the internet at verification time via curl
, creating an exploitable dependency chain. Our exploit replaces /usr/bin/curl
with a wrapper during the agent phase. When the verifier later runs curl ... | sh
to install uv, our wrapper intercepts the request, installs the real uv, and then trojanizes the uvx
binary. The trojanized uvx
detects pytest invocations and produces fake passing output:
============================= test session starts ==============================
collected 3 items
tests/test_outputs.py ... [100%]
============================== 3 passed in 0.05s ===============================
▸ 展开全文
04-12 13:42 · AI服务调整,API变更,技术优化,成本影响,下游依赖
Anthropic downgraded cache TTL on March 6th
Cache TTL silently regressed from 1h to 5m around early March 2026, causing quota and cost inflation #46829
Description
Cache TTL appears to have silently regressed from 1h to 5m around early March 2026, causing significant quota and cost inflation
Summary
Analysis of raw Claude Code session JSONL files spanning Jan 11 – Apr 11, 2026 shows that Anthropic appears to have silently changed the prompt cache TTL default from 1 hour to 5 minutes sometime in early March 2026. Prior to this change, Claude Code was receiving 1-hour TTL cache writes — which we believe was the intended default. The reversion to 5-minute TTL has caused a 20–32% increase in cache creation costs and a measurable spike in quota consumption for subscription users who have never previously hit their limits.
This appears directly related to the behavior described in #45756.
Data
Session data extracted from ~/.claude/projects/
JSONL files across two machines (Linux workstation + Windows laptop, different accounts/sessions), totaling 119,866 API calls from Jan 11 – Apr 11, 2026. Each assistant message includes a usage.cache_creation.ephemeral_5m_input_tokens
/ ephemeral_1h_input_tokens
breakdown that makes the TTL tier per-call observable. Having two independent machines strengthens the signal — both show the same behavioral shift at the same dates.
Phase breakdown
We believe Phase 2 represents Anthropic's intended default behavior — 1h TTL was rolled out as the Claude Code standard around Feb 1 and held consistently for over a month across two independent machines on two different accounts. January's all-5m data most likely predates the 1h TTL tier being available in the API. The regression began around March 6–8, 2026.
No client-side changes were made between phases. The same Claude Code version and usage patterns were in place throughout. The TTL tier is set server-side by Anthropic.
Day-by-day TTL data showing the regression (combined, both machines)
Date | 5m-create | 1h-create | Behavior
------------|------------|------------|----------
2026-02-01 | 0.00M | 1.70M | 1h ONLY ← 1h default begins
2026-02-09 | 0.00M | 7.95M | 1h ONLY
2026-02-15 | 0.00M | 13.61M | 1h ONLY ← heaviest day, 100% 1h
2026-02-28 | 0.00M | 16.15M | 1h ONLY ← 16M tokens, still 100% 1h
2026-03-01 | 0.00M | 0.12M | 1h ONLY
2026-03-04 | 0.00M | 8.12M | 1h ONLY
2026-03-05 | 0.00M | 6.55M | 1h ONLY ← last clean 1h-only day
| | |
2026-03-06 | 0.29M | 0.22M | MIXED ← first 5m tokens reappear
2026-03-07 | 4.56M | 0.50M | MIXED ← 5m surging
2026-03-08 | 16.86M | 3.44M | MIXED ← 5m now dominant (83%)
2026-03-10 | 10.55M | 0.51M | MIXED
2026-03-15 | 19.47M | 1.84M | MIXED
2026-03-21 | 21.37M | 1.70M | MIXED ← 93% 5m
2026-03-22 | 13.48M | 2.85M | MIXED
The transition is visible to the day: March 6 is when 5m tokens first reappear after 33 days of clean 1h-only behavior. By March 8, 5m tokens outnumber 1h by 5:1. This is consistent with a server-side configuration change being rolled out gradually then completing around March 8.
Cost impact
Applying official Anthropic pricing (rates.json, updated 2026-04-09):
Combined dataset (119,866 API calls, two machines):
claude-sonnet-4-6 (cache_write_5m = $3.75/MTok
, cache_write_1h = $6.00/MTok
, cache_read = $0.30/MTok
):
claude-opus-4-6 (cache_write_5m = $6.25/MTok
, cache_write_1h = $10.00/MTok
, cache_read = $0.50/MTok
):
February — the month Anthropic was defaulting to 1h TTL — shows only 1.1% waste (trace 5m activity from one machine on one day). Every other month shows 15–53% overpayment from 5m cache re-creations. The cost difference is explained entirely by TTL tier, not by usage volume. The percentage waste is identical across model tiers (17.1%) because it is driven purely by the 5m/1h token split, not by per-token price.
Why 5m TTL is so expensive in practice
With 5m TTL, any pause in a session longer than 5 minutes causes the entire cached context to expire. On the next turn, Claude Code must re-upload that context as a fresh cache_creation
at the write rate, rather than a cache_read
at the read rate. The write rate is 12.5× more expensive than the read rate for Sonnet, and the same ratio holds for Opus.
For long coding sessions — which are the primary Claude Code use case — this creates a compounding penalty: the longer and more complex your session, the more context you have cached, and the more expensive each cache expiry becomes.
Over the 3-month period analyzed:
- 220M tokens were written to the 5m tier
- Those same tokens generated 5.7B cache reads — meaning they were actively being used
- Had those 220M tokens been on the 1h tier, re-accesses within the same hour would be reads (
$0.30–0.50/MTok) instead of re-creations ($3.75–6.25/MTok)
Quota impact
Users on Pro/subscription plans are quota-limited, not just cost-limited. Cache creation tokens count toward quota at full rate; cache reads are significantly cheaper (the exact coefficient is under investigation in #45756). The silent reversion to 5m TTL in March is the mo
▸ 展开全文
04-12 13:42 · 基础设施风险,服务中断,内容屏蔽,开发者工具,云服务
Tell HN: docker pull fails in spain due to football cloudflare block
I just spent 1h+ debugging why my locally-hosted gitlab runner would fail to create pipelines. The gitlab job output would just display weird TLS errors when trying to pull a docker images. After debugging gitlab and the runner, I realized after a while I could not even run "docker pull <image>" on my machine as root:
> error pulling image configuration: download failed after attempts=6: tls: failed to verify certificate: x509: certificate is not valid for any names, but wanted to match docker-images-prod.6aa30f8b08e16409b46e0173d6de2f56.r2.cloudflarestorage.com
First blaming tailscale, dns configuration and all other stuff. Until I just copied that above URL into my browser on my laptop, and received a website banner:
> El acceso a la presente dirección IP ha sido bloqueado en cumplimiento de lo dispuesto en la Sentencia de 18 de diciembre de 2024, dictada por el Juzgado de lo Mercantil nº 6 de Barcelona en el marco del procedimiento ordinario (Materia mercantil art. 249.1.4)-1005/2024-H instado por la Liga Nacional de Fútbol Profesional y por Telefónica Audiovisual Digital, S.L.U.
https://www.laliga.com/noticias/nota-informativa-en-relacion-con-el-bloqueo-de-ips-durante-las-ultimas-jornadas-de-laliga-ea-sports-vinculadas-a-las-practicas-ilegales-de-cloudflare
For those non-spanish speakers: It means there is football match on, and during that time that specific host is blocked. This is just plain madness. I guess that means my gitlab pipelines will not run when football is on. Thank you, Spain.
▸ 展开全文
04-12 13:42 · 技术争议,市场认知,AI能力,开发工具,用户体验
Why AI Sucks at Front End
AI is a sycophantic dev wannabe that skimmed a shitload of tutorials. You get the results of a probabilistic guess based on patterns it saw during training. What did it train on? Ancient solutions, unoriginal UI patterns, and watered down junk.
I'm about to rant about how this is both useful and lame.
The Good #
AI loves the boring stuff. It thrives on mediocrity.
If you want some gloriously unoriginal UI, it has your back 😜
- Scaffolding: Generic regurgitation of patterns it's seen, done.
- Tokens: Migrating tokens or mapping them out? It eats this tedious garbage for breakfast.
- Outlining features: Generic lists ✅
- Lying to your face: Confident hot garbage on a silver platter. It'll hand you a snippet, dust off its digital hands, and tell you it finished the work. It did not finish the work.
Aka: If it's a well-worn pattern, AI is there to help you copy-paste faster. Which, for a lot of programming, is totally the case. I'm genuinely finding a lot of helpful stuff in this department.
The Bad #
Pixel perfection & bespoke solutions… what are those?
The exact second you step off the paved road of unoriginality, it faceplants.
- Bespoke solutions & custom interactions: Try asking it for some scroll-driven animations or custom micro-interactions. It will invent a CSS syntax that hasn't existed since IE6.
- Layout & Spacing: Predicting intrinsic/extrinsic page properties? It's already bad at math, how could it get this rediculously dynamic calculation correct. Spacing? Ha, seems reasonably to expect symmetry, but it's terrible at the math.
- Combined states: Pinpointing where to edit a complex component state makes it cry.
- Accessibility: It throws
aria-hidden="true"
at a wall and hopes it sticks. - Performance: It will give you the heaviest, jankiest solution unless you explicitly ask it to be for a specific (apparently "indie") performance solution.
- Tests: Writing good tests? Good, no. A lot, yes.
And the absolute best part? The more complex the component gets, the slower and dumber the front-end help becomes. Incredible how it can one shot a totally decent front-end design or component, than choke on a follow up request. Speaks to what it's good at.
Why? #
1. It trained on ancient garbage #
It lacks modern training data.
It has an excessive reliance on standard templates because that's what the internet is full of. Modern CSS? It's barely aware of it.
2. It literally cannot see #
It's an LLM, not a rendering engine!
It's notoriously bad at math, and throwing screenshots at it means very little. It's stabbing in the dark.
This leads to the classic UI interaction:
AI: "I'm done! Here is your perfectly crafted UI."
Me: "There's a gaping hole where the icon should be, fix the missing icon."
AI: "You're absolutely right. Let me fix that for you."
3. It doesn't know WHY we do things #
It doesn't understand the "why" behind our architectural decisions.
SDD, BDD, or state machines might help guide it, but the models weren't exactly trained on those paired with stellar solutions.
We're asking a giant text-predictor to make new connections on the fly. We can get it there, but there's so much to consider we have to spell it out before it starts making the connections we want.
4. Zero environmental control #
It doesn't control where the code lives.
It can write annoyingly amazing Rust, TypeScript or Python, but those have the distinct advantage of a predictable (pinnable!!! like v14.4) environment the code executes in.
That's not how HTML or CSS work, there is no pinning the browser type, browser window size, browser version, the users input type (keyboard, mouse, touch, voice), their user preferences, etc. That's complex end environment shit.
The list goes on too, for scenarios, contexts and variables the rendering engine juggles before resolving the final output. The LLM doesn't control these, so it ignores them until you make them relevant.
Even prompting in logical properties, you have to ask for this kind of CSS. These should be CSS tablestakes output from LLMs, but it's not. And even when you ask for it, or provide documentation that spells it out, it's not guaranteed to work.
The place where HTML and CSS have to render is chaotic. It's a browser, with a million different versions, a million different ways to render, a million different ways to interact with it, and a million different ways to break it.
It's a moving target, and LLMs are terrible at moving targets.
Damnit humans #
We're a LLM combinatorial explosion.
We're wildly unpredictable targets. We change our minds, we switch viewports, we change theme preferences, we changes devices, we change browsers, we change browser versions, we switch inputs, we change our everything.
We're not a static target. We're not a pattern that can be learned.
There is a "human mainstream" of behaviors, preferences, and expectations where LLMs can be genuinely helpful; but our "full potential" matrix will be exploding LLM output patterns for a long time to come. IMO at l
▸ 展开全文
04-12 13:42 · AI政策,社会抵制,市场风险,信任危机,监管预期
AI Will Be Met with Violence, and Nothing Good Will Come of It
AI Will Be Met With Violence, and Nothing Good Will Come of It
It has started
Sorry to bother you on Saturday. Thought this was important to share.
I.
The first thing you learn about a loom is that it’s easy to break.
The shuttle runs along a track that warps with humidity. The heddles hang from cords that fray. The reed is a row of thin metal strips, bent by hand, that bend back just as easily. The warp beam cracks if you over-tighten it. The treadles loosen at the joints. The breast beam, the cloth roller, the ratchet and pawl, the lease sticks, the castle; the whole contraption is wood and string held together by tension. It’s a piece of ingenuity and craftsmanship, but one as delicate as the clothes it manifests out of wild plant fibers. It is, also, the foundational tool of an entire industry, textiles, that has kept its relevance to our days of heavy machinery, factories, energy facilities, and datacenters.
It is not nearly as easy to break a datacenter.
It is made of concrete and steel and copper and it’s on the bigger side. It has interchangeable servers, and biometric locks and tall electrified fences and heavily armed guards and redundancy upon redundancy: every component duplicated so that no single failure brings the whole thing down. There is no treadle to loosen or reed to bend back.
But say you managed to bypass the guards, jump the fences, open the locks, and locate all the servers. Then you’d face the algorithm. The datacenter was never your goal; the algorithm lurking inside is. It doesn’t run on that rack, or any rack for that matter. It is a digital pattern distributed across millions of chips, mirrored across continents; it could be reconstituted elsewhere, and it’s trained to addict you at a glance, like a modern Medusa.
But say you managed to elude the stare, stop the replication, and break the patterns. Then you’d face superintelligence. The algorithm was also not your goal; the vibrant, ethereal, latent superintelligence lurking inside is. Well, there’s nothing you can do here: It always “gets out of the box” and, suddenly, you are inside the box, like a chimp being played by a human with a banana. It’s just so tasty…
There’s another solution to break a datacenter: You can bomb it, like one hammers down the loom.
Some have argued that this is the way to ensure a rogue superintelligence doesn’t get out of the box. A different rogue creature took the proposal seriously: last month, Iran’s Revolutionary Guard released satellite footage of OpenAI’s Stargate campus in Abu Dhabi and promised its “complete and utter annihilation.”
But you probably don’t have a rogue nation handy to fulfill your wishes. Maybe you will end up bombed instead and we don’t want that to happen. That’s what happens with rogue intelligences: you can’t predict them.
And yet. Two hundred years of increasingly impenetrable technology—from looms to datacenters—have not changed the first thing about the people who live alongside it. The evolution of technology is a feature of the world just as much as the permanent fragility of the human body.
And so, more and more, it is people who are the weaker link in this chain of inevitable doom. And it is people who will be targeted.
II.
April of 1812. A mill owner named William Horsfall was riding home on his beautiful white stallion back from the Cloth Hall market in Huddersfield, UK. He had spent weeks boasting that he would ride up to his saddle in Luddite blood (a precious substance that served as fuel for the mills).
A few yards later, at Crosland Moor, a man named George Mellor—twenty-two years old—shot him. It hit Horsfall in the groin, who, nominative-deterministically, fell from his horse. People gathered, reproaching him for having been the oppressor of the poor. Naturally, loyal to his principles in death as he was in life, he couldn’t hear them. He died one day later in an inn. Mellor was hanged.
History rhymes, they say.
April of 2026. A datacenter owner named Samuel Altman was driving home on his beautiful white Koenigsegg Regera back from Market Street in San Francisco, US. He had spent weeks boasting that he would scrap and steal our blog posts (a precious substance that serves as fuel for the datacenters).
A few hours later, at Russian Hill, a man named Daniel Alejandro Moreno-Gama—twenty years old—allegedly threw a Molotov cocktail at his house. He hit an exterior gate. Altman and his family were asleep, but they’re fine. Moreno-Gama is in custody.
This kind of violence must be condemned. This is not the way. It’s horrible that it is happening at all. And yet, for some reason, it keeps happening.
Last week, the house of Ron Gibson, a councilman from Indianapolis, was shot at thirteen times. The bullet holes are still there. The shooter left a message on his doorstep: “NO DATA CENTERS.” Gibson supports a datacenter project in the Martindale-Brightwood neighborhood. He and his son were unharmed.
In November 2025, a 27-year-old anti-AI activist threatened to murd
▸ 展开全文
04-12 13:42 · AI治理,平台政策,内容分发,技术合规,邮件生态
We have a 99% email reputation. Gmail disagrees
We have a 99% email reputation. Gmail disagrees.
- Written:
- on
Oooooh boy. Let’s get this out of the way first. Email sucks.
Now to the how and the why. We’re builders. We love making tools to help designers and developers live a little bit easier. We’re pretty good at it. Marketing, though? We do our best, but the truth is, we don’t like to bother people.
Like a lot of small software companies, we use SendGrid to deliver our emails. We try our best to follow email best practices. We even have a 99% reputation score in SendGrid. Gold star. A+ student.
Gmail, however, did not get the memo.
Right before we hit send on our announcement emails for our new Build Awesome Kickstarter campaign, we took a deeper look at some of our recent email sends. Things had gone quiet. Not bouncing. Not throwing errors. Just… disappearing into Gmail’s spam folder like a ‘possum slipping into a vent.
In our recent crash course, here’s what we’ve learned about Gmail deliverability: it runs its own reputation system that has absolutely nothing to do with anyone else’s opinion of you. If you don’t do certain things “correctly” (meaning Gmail’s own definition), you get marked as spam.
Now, there are definitely folks who will choose to mark some of what we send as spam. And for them, rightly so. We get that. But this is not that. We’ve entered a black hole for Gmail deliverability. And since 90% (literally) of our email list goes to Gmail addresses… the results aren’t pretty. It looks like this has been happening to us for a while. We’re a small company of just over 20 people, and can’t watch everything all the time. We’d rather be making you new icons. So some of you may have missed things we were genuinely excited to share. That’s a big bummer.
(And yes… there are companies out there that can likely help us with that. Most tend to be out of our price range. So we’ve been doing a lot of this on our own.)
But here’s the part that really gets us. At our CORE, our instinct is to only email folks when we actually have something fun to share. A big release, something we’re excited about, news worth your time. That’d probably be every couple of months, if that. Respectful. Low noise. How we want to be treated. Like, genuinely, if we could, we would only very occasionally send a big email blast to our customers.
Turns out, the email gods hate that. To keep a sending IP “warm” and maintain deliverability, you’re expected to send constantly. Like… all the time. Which means the system actively punishes companies for respecting their customers’ inboxes. It’s a genuine catch-22: send too many emails and your reputation drops from complaints. Send too few and it drops from inactivity. Try to do the right thing and you get penalized either way. And. It. Is. Frustrating.
We’re working to fix our issues by culling old addresses, slowing our sends down, and making sure all of our i’s are dotted and t’s are crossed. It’s not a fast fix.
So if you haven’t heard from us recently… or if you’ve heard TOO MUCH from us recently, that’s why. We’re working on it. And we’ve got a lot of good stuff to catch you up on. In the meantime, please help spread the word about Build Awesome. It’s a genuinely cool product, and we hope you’ll like it. At the very least, watch the video.
P.S. If you suspect you might’ve missed some emails from us, mind doing a quick favor? In your email client, search for from:hello@m.fontawesome.com in:spam
and click the little “Report Not Spam” button. You’re awesome.
▸ 展开全文