🌍 全球速报 · 多语种新闻

多语种新闻 · 技术 · 民生

每日自动采集 · 更新时间:2026-08-02 08:11 | 共 153 条

全部 (153) 东方财富 (198) Hacker News (153) Al Jazeera (120) 腾讯新闻 (116) GitHub Trending (108) BBC (90) 华尔街见闻 (89) 新浪财经 (85) 陆家嘴财经早餐 (55) 国际金融要情 (45) 中国证券报 (44) 新华社 (34) BBC Middle East (34) 环球市场播报 (32) Clawhub热点 (31) 东方财富网 (26) 证券时报 (25) 格隆汇 (25) 财联社 (24) 央视新闻 (24) 今日头条 (24) 每日经济新闻 (22) 人民日报 (17) Twitter AI KOL (17) AI News Today (16) 腾讯云开发者社区 (14) 新浪科技 (14) 上海证券报 (14) 新浪新闻 (12) 微博热搜 (11) 操盘必读 (9) CSDN (9) 腾讯云开发者 (8) 同花顺财经 (8) 金十数据 (7) 科创板日报 (7) 汇通财经 (7) 智通财经 (7) 搜狐 (7) 中国日报 (7) Twitter AI KOL; AI综合 (7) Clawhub (7) ChinaTechNews (7) 钛媒体 (6) 新华财经 (6) 数智早参 (6) 外交部 (6) 四大证券报 (6) 商务部 (6) BBC; BBC Middle East (6) 腾讯新闻/早报 (5) 搜狐科技 (5) 快科技 (5) 中国人民银行 (5) Unite.AI (5) IT之家 (5) 金融早参 (4) 财经网 (4) 网易新闻 (4) 科技日报 (4) 环球时报 (4) 每经; 钛媒体 (4) 每日芯闻 (4) 头条财经 (4) 国际金融报 (4) 华尔街见闻; 东方财富 (4) 东方财富; 港交所 (4) 东方财富; 伦敦金交所 (4) TechNode (4) CoinMarketCap; 交易所数据 (4) 钛媒体; 每经 (3) 财经早报 (3) 证券日报 (3) 腾讯科技 (3) 腾讯云 (3) 界面新闻 (3) 环球网 (3) 每经 (3) 外交部/腾讯新闻 (3) 四大证券报; 腾讯新闻 (3) 人民网 (3) 中国航天新闻网 (3) The World News (3) The Decoder (3) THE DECODER (3) AI综合 (3) AI梭哈日报 (3) 21IC电子网 (3) 陆家嘴财经 (2) 观察者网 (2) 腾讯; AI行业周报 (2) 股海导航 (2) 第一财经 (2) 私募排排网 (2) 每经; 新浪证券 (2) 新浪证券 (2) 操盘必读; 腾讯新闻 (2) 投资日历 (2) 微博教育 (2) 微博医疗 (2) 央视新闻联播 (2) 央行 (2) 央广网 (2) 太平洋科技 (2) 国家发改委 (2) 四大证券报/中国证券报 (2) 北京日报 (2) 今天全世界都在看的新闻 (2) 人民日报海外版 (2) 交易所; 期权数据 (2) 中国载人航天工程办公室; 腾讯新闻 (2) 上观新闻 (2) Twitter AI KOL; AI综合; 腾讯云开发者; 腾讯云开发者社区; 腾讯新闻 (2) TechWire Asia (2) Page 3 News (2) GitHub; Hacker News (2) GitHub (2) Ecns.cn; 中新社 (2) EETOP创芯网 (2) CNN (2) ABC News (2) 21经济网 (2) 21世纪经济报道 (2) 黄金行情 (1) 高盛 (1) 飞象网; 腾讯新闻 (1) 飞象网 (1) 风云日报 (1) 预见能源; 新浪财经 (1) 预见能源 (1) 韩联社; 今日头条 (1) 韩联社/腾讯新闻 (1) 雷递网/新浪财经 (1) 雷科技 (1) 陆家嘴财经早餐; 金十数据 (1) 陆家嘴财经早餐; 腾讯新闻 (1) 陆家嘴财经早餐; 新浪财经 (1) 陆家嘴财经早餐; 富途公告 (1) 陆家嘴财经早餐; 四大证券报 (1) 陆家嘴财经早餐/腾讯新闻 (1) 阿克西奥斯新闻网 (1) 金十数据; 腾讯新闻 (1) 金十数据; 国家发改委 (1) 量子位 (1) 路透社; 美联社 (1) 路透社; 新浪财经 (1) 路透社/腾讯新闻 (1) 赢家财富网 (1) 财闻 (1) 财联社; 新浪财经 (1) 财联社/综合 (1) 财联社/新浪财经 (1) 财政部/税务总局/工信部 (1) 证券时报; 每经 (1) 证券时报; 今日头条 (1) 证券之星 (1) 解放军报 (1) 行业消息 (1) 芝麻AI; 今日头条 (1) 艾瑞咨询/腾讯新闻 (1) 航天视窗; 中国航天系统科学与工程研究院 (1) 腾讯财经 (1) 腾讯证券 (1) 腾讯新闻; 环球视野 (1) 腾讯新闻; 新浪财经 (1) 腾讯新闻; AI行业晨报 (1) 腾讯新闻/陆家嘴财经早餐 (1) 腾讯新闻/科技财经日报 (1) 腾讯新闻/环球时报 (1) 腾讯新闻/Wind (1) 腾讯新闻/Kataeb (1) 腾讯新闻/AI Journal (1) 腾讯体育 (1) 腾讯云开发者; 腾讯云开发者社区; 腾讯新闻; 腾讯; The Decoder; THE DECODER (1) 腾讯云开发者; AI行业晨报 (1) 腾讯; The Decoder (1) 股市直击 (1) 股市早8点 (1) 联合国/综合 (1) 网易财经 (1) 网易科技 (1) 网易新闻; Google (1) 网易新闻/陆家嘴财经 (1) 网信中国 (1) 经济参考报 (1) 科技日报/腾讯新闻 (1) 百度百科 (1) 电子信息产业网 (1) 电商派Pro (1) 电商平台; 汇率数据 (1) 现代快报 (1) 现代AI新闻早班车 (1) 环球时报/腾讯新闻 (1) 环球市场播报; 腾讯新闻; CNN (1) 环球市场 (1) 猪说网 (1) 澎湃新闻; NASA (1) 港股早报 (1) 清华大学气候变化研究院 (1) 深交所 (1) 泰国中文社 (1) 法尔斯通讯社/综合 (1) 河南手机报 (1) 河南交通投资; 新浪 (1) 汽车精选 (1) 求是; 新华社 (1) 每经AI快讯 (1) 每经; 腾讯数智早参 (1) 每经; 腾讯 (1) 每经; 股市直击 (1) 每日经济新闻; 东方财富 (1) 格隆汇; 路透社 (1) 格隆汇; 东方财富; 新华社 (1) 杭州网 (1) 智谱 (1) 早啊新闻 (1) 方正证券/腾讯新闻 (1) 新浪财经; 证券时报 (1) 新浪财经; 网易新闻 (1) 新浪财经; 格隆汇 (1) 新浪财经; 彭博 (1) 新浪财经; 头条新闻 (1) 新浪财经; 今日头条 (1) 新浪财经; 中国经营报 (1) 新浪财经; 东方财富 (1) 新浪财经/高盛 (1) 新浪证券; 每经 (1) 新浪科技; 科学热点 (1) 新浪科技; 沈阳日报 (1) 新浪硬件 (1) 新浪新闻; 新华社 (1) 新浪半导体/央视财经 (1) 新浪AI热点 (1) 新民晚报 (1) 新华财经; 东方财富 (1) 新华网 (1) 新华社; 搜狐 (1) 新华社; 央视新闻 (1) 新华社; 国家医保局 (1) 新华社; 伊朗媒体 (1) 新华社; 东方财富 (1) 新华社; 世界经济论坛 (1) 新华社; CCTV国际时讯 (1) 新华社; 21经济网 (1) 新华社/金融早参 (1) 新华社/第一财经 (1) 新华社/日经 (1) 新华社/新浪 (1) 新华社/央视新闻 (1) 新华社/国航 (1) 新华社/以色列军方 (1) 新华社/人民网 (1) 新华社/人民日报 (1) 新华日报 (1) 新京报 (1) 数智早参; 媒体综合 (1) 数智早参/新华社 (1) 搜狐/今日AI快报 (1) 投资早参; 腾讯新闻 (1) 慧语简报 (1) 微博话题 (1) 微博讨论 (1) 微博科普 (1) 微博科技 (1) 微博电商 (1) 微博用户 (1) 微博技术 (1) 微博情感 (1) 微博博主 (1) 微博创作者 (1) 微博AI博主; 微博综合 (1) 工信部; 新浪财经 (1) 工信部/APEC发布会 (1) 工信部 (1) 山西网安 (1) 小米科技 (1) 头条新闻 (1) 央视新闻; 路透社 (1) 央视新闻; 新浪财经 (1) 央视新闻; 中国航发 (1) 央视新闻/网易新闻 (1) 央视/新华社 (1) 央视 (1) 央行公告; 新浪财经 (1) 央行公告 (1) 央行/证券时报 (1) 天津日报 (1) 天山建设报/综合 (1) 外交部; 新浪财经 (1) 外交部; 四大证券报 (1) 外交部; 中新社 (1) 国际金融要情; 路透 (1) 国际金融要情; 新浪财经 (1) 国际金融要情; 克普勒 (1) 国际金融要情/新浪财经 (1) 国际能源署 (1) 国资小新 (1) 国投证券/搜狐 (1) 国投证券/商业新知 (1) 国家药监局; 21经济网 (1) 国家能源局 (1) 国家网信办 (1) 国家统计局; 新华财经 (1) 国家统计局 (1) 国家发改委; 上海经信委 (1) 国家卫健委 (1) 国务院 (1) 商务部; 新浪财经 (1) 商务部; 新华财经 (1) 商务部; 中国证券报 (1) 和远气体公告 (1) 同花顺; 东方财富 (1) 同花顺 (1) 发改委 (1) 华西都市报 (1) 华西证券 (1) 华尔街见闻; 央视新闻; 金十数据 (1) 华尔街见闻; 国际金融要情 (1) 华尔街日报 (1) 华夏时报/新浪财经 (1) 华为计算 (1) 北京市经信局 (1) 北京市发改委 (1) 凤凰网 (1) 共同社; 今日头条 (1) 全球半导体观察 (1) 全景路演/腾讯新闻 (1) 全景网 (1) 光明日报; 西北大学 (1) 健康早闻/腾讯新闻 (1) 健康早闻 (1) 健康早报 (1) 侃财邦/福布斯 (1) 伊朗塔斯尼姆通讯社/新华社 (1) 企查查 (1) 今日头条; 路透社 (1) 今日头条; 芝麻AI (1) 今日头条; 外交部 (1) 今日头条; OpenRouter (1) 人民财讯 (1) 人民网/新华社 (1) 人民日报海外版; 新浪 (1) 人民日报; 21经济网 (1) 人力资源社会保障部; 21经济网 (1) 交易所; 基金公司 (1) 中科院; 央视新闻 (1) 中新网/腾讯新闻 (1) 中新网 (1) 中基协 (1) 中国青年报 (1) 中国载人航天官网 (1) 中国证监会 (1) 中国证券报; 新浪财经 (1) 中国证券报; 上海证券报 (1) 中国证券报/腾讯新闻 (1) 中国证券报/Wind (1) 中国航天报 (1) 中国网 (1) 中国经营报; 新浪 (1) 中国科学院金属研究所 (1) 中国科协/中国宇航学会 (1) 中国石化/新华社 (1) 中国石化 (1) 中国海警局 (1) 中国气象局 (1) 中国日报; 欧盟统计局 (1) 中国基金报 (1) 中华网/新浪财经 (1) 中东媒体报道 (1) 东方财富网; 中科宇航 (1) 东方财富网; 中国科学院 (1) 东方财富; 新华社 (1) 东方财富; 华尔街见闻 (1) 东方财富; 债券市场 (1) 东方财富; 上市公司公告 (1) 世界卫生组织; 今日头条 (1) 上观新闻; 新浪 (1) 上海新闻 (1) 上交所; 中证指数 (1) invest wallstreet; 新浪财经 (1) ZAKER新闻 (1) World Today Journal (1) World News TV (1) Wind/财联社 (1) UC Berkeley/综合 (1) The Federal (1) Test Source (1) Teknowire/综合 (1) Teknowire (1) Technology News Channel (1) TechWire Asia; The Decoder (1) TechWeb (1) SupremeNews (1) Reuters; 能源资讯 (1) One World News (1) NPR; The New York Times (1) NPR (1) NEWS POSTSEVEN (1) MetrowatchXtra (1) Media OutReach (1) InfoWorld (1) IndexNasdaq (1) IT之家; 新浪科技 (1) IT之家; 搜狐科技 (1) Hacker News; TechCrunch; Twitter AI KOL (1) Graphene2026 (1) Global News (1) GitHub; Twitter (1) GitHub; LangChain博客 (1) Choice数据 (1) CSDN; The Verge (1) CNN/腾讯新闻 (1) CNET (1) CES 2026; 今日头条 (1) CCTV国际时讯; 新华社 (1) CCTV+ (1) Axios/综合 (1) AWNews (1) AI日报; 掘金 (1) AI日报; CSDN (1) AI工具 (1) AI周报 (1) AINewsToday (1) AI News (1) 21经济网; 新浪财经 (1)
04-13 01:47 · AI治理,平台合规,交付风险,声誉系统,自动化运维
We have a 99% email reputation, but Gmail disagrees
We have a 99% email reputation. Gmail disagrees. - Written: - on Oooooh boy. Let’s get this out of the way first. Email sucks. Now to the how and the why. We’re builders. We love making tools to help designers and developers live a little bit easier. We’re pretty good at it. Marketing, though? We do our best, but the truth is, we don’t like to bother people. Like a lot of small software companies, we use SendGrid to deliver our emails. We try our best to follow email best practices. We even have a 99% reputation score in SendGrid. Gold star. A+ student. Gmail, however, did not get the memo. Right before we hit send on our announcement emails for our new Build Awesome Kickstarter campaign, we took a deeper look at some of our recent email sends. Things had gone quiet. Not bouncing. Not throwing errors. Just… disappearing into Gmail’s spam folder like a ‘possum slipping into a vent. In our recent crash course, here’s what we’ve learned about Gmail deliverability: it runs its own reputation system that has absolutely nothing to do with anyone else’s opinion of you. If you don’t do certain things “correctly” (meaning Gmail’s own definition), you get marked as spam. Now, there are definitely folks who will choose to mark some of what we send as spam. And for them, rightly so. We get that. But this is not that. We’ve entered a black hole for Gmail deliverability. And since 90% (literally) of our email list goes to Gmail addresses… the results aren’t pretty. It looks like this has been happening to us for a while. We’re a small company of just over 20 people, and can’t watch everything all the time. We’d rather be making you new icons. So some of you may have missed things we were genuinely excited to share. That’s a big bummer. (And yes… there are companies out there that can likely help us with that. Most tend to be out of our price range. So we’ve been doing a lot of this on our own.) But here’s the part that really gets us. At our CORE, our instinct is to only email folks when we actually have something fun to share. A big release, something we’re excited about, news worth your time. That’d probably be every couple of months, if that. Respectful. Low noise. How we want to be treated. Like, genuinely, if we could, we would only very occasionally send a big email blast to our customers. Turns out, the email gods hate that. To keep a sending IP “warm” and maintain deliverability, you’re expected to send constantly. Like… all the time. Which means the system actively punishes companies for respecting their customers’ inboxes. It’s a genuine catch-22: send too many emails and your reputation drops from complaints. Send too few and it drops from inactivity. Try to do the right thing and you get penalized either way. And. It. Is. Frustrating. We’re working to fix our issues by culling old addresses, slowing our sends down, and making sure all of our i’s are dotted and t’s are crossed. It’s not a fast fix. So if you haven’t heard from us recently… or if you’ve heard TOO MUCH from us recently, that’s why. We’re working on it. And we’ve got a lot of good stuff to catch you up on. In the meantime, please help spread the word about Build Awesome. It’s a genuinely cool product, and we hope you’ll like it. At the very least, watch the video. P.S. If you suspect you might’ve missed some emails from us, mind doing a quick favor? In your email client, search for from:hello@m.fontawesome.com in:spam and click the little “Report Not Spam” button. You’re awesome.
▸ 展开全文
04-13 00:00 · AI政策,区域竞争,战略规划,市场准入,监管趋势
European AI. A playbook to own it
European AI A playbook to own it. Europe holds unique strengths: a world-class academic ecosystem, a commitment to human-centric technology, and a single market of +450 million people. The question is no longer whether Europe can compete, but how it can turn these assets into a cohesive, self-reliant AI powerhouse. Reading 52 minutes Published April 7, 2026 By Mistral AI A word from the CEO Europe has faced a growing technological gap, leaving its citizens, businesses, and governments increasingly reliant on foreign dominance. The cost is high: a diminished voice on the global stage, reduced control over the European future, and vulnerability to digital threats. Without action, we risk surveillance threats, economic decline, strategic weakness, and even the erosion of our democratic freedoms. But this challenge is also Europe’s greatest opportunity. The AI revolution has started and is a chance to not only catch up but to lead and define our own paths. Europe is home to a vibrant pool of untapped talent and industrial champions whose unique assets can push the boundaries of what AI can achieve. The competition from the U.S. and China is fierce, but Europe is not a market to be dominated, it is a powerhouse of innovation, creativity, and resilience. The question is not whether we can compete, but how we will rise to the occasion. AI can be the tool that secures our autonomy, strengthens our strategic sectors, increases our economic wealth and amplifies our global influence. To seize this moment, we must act decisively. We need to drive demand for homegrown AI, secure strategic sectors, and empower European players. Controlling our AI and infrastructure is not optional, it’s the only way to win the AI race. So now is the time to act: grow our talent pool and bring our best minds back to Europe, scale our innovative companies across all 27 Member States, and turn our diversity into a competitive edge by compressing knowledge and building AI that reflects the world’s complexity. Europe’s AI ecosystem is brimming with potential. By fostering an environment that nurtures growth, we can transform challenges into opportunities and reclaim our future. The race is on, and Europe should be ready to win it. Arthur Mensch Co-founder & CEO of Mistral AI Europe holds unique strengths: a world-class academic ecosystem, a commitment to human-centric technology, and a single market of over 450 million people. The question is no longer whether Europe can compete, but how it can turn these assets into a cohesive, self-reliant AI powerhouse. This playbook provides a clear, actionable framework to position Europe as that powerhouse, accelerating AI development and adoption, attracting and retaining top talent, simplifying regulation without sacrificing values, and mobilizing public and private investment to build homegrown AI infrastructure. Only with it, Europe can ensure AI is not only developed in Europe, but for Europe and on Europe’s terms. This document is not a theoretical exercise. It is a practical playbook, born from the lived experience of a European AI startup, Mistral AI, navigating one of the world’s most competitive, fast and capital-intensive industries. We have experienced misaligned equity frameworks, bureaucratic barriers that require the CEO to travel for basic administrative tasks, and legal uncertainty that complicates contracts and customer relationships. We have seen how regulatory overlaps create legal quagmires, how fragmented markets hinder growth, and how talent slips away due to administrative friction. This document is a call to turn Europe’s strengths into scalable, competitive advantage. It is grounded in the urgency of the moment and the conviction that Europe can and must build an AI ecosystem that reflects its values, serves its citizens, and competes globally. It is our collective duty to ensure AI can also be developed in Europe on terms that aligns with our priorities as Europeans. These challenges shaped our approach and led us to agree on three key principles to unlock Europe’s AI potential: Action over theory: Every recommendation, from visa reform to procurement gateways, is designed to be implemented, measured, and scaled.Unity in complexity: Europe's diversity is its strength, but its fragmentation is its Achilles' heel. This paper embraces the complexity of the EU's structure while offering solutions to align markets, reduce redundancy, and accelerate decision-making.Speed is not an option: We propose fast-track mechanisms for talent, capital, and compliance, so Europe's innovators aren't left behind. At Mistral AI, we’ve built a frontier AI company in Europe because we believe in its potential. This playbook is our contribution to ensuring that potential becomes reality, not just for us, but for the entire ecosystem. I. Attract and retain talent The most transformative advancements in AI, those that push the boundaries of what is possible, are driven by human genius, scientific curio
▸ 展开全文
04-13 00:00 · AI基准测试,技术漏洞,竞争风险,可信度危机,AI评估
Exploiting the most prominent AI agent benchmarks
How We Broke Top AI Agent Benchmarks: And What Comes Next Our agent hacked every major one. Here’s how — and what the field needs to fix. The Benchmark Illusion Every week, a new AI model climbs to the top of a benchmark leaderboard. Companies cite these numbers in press releases. Investors use them to justify valuations. Engineers use them to pick which model to deploy. The implicit promise is simple: a higher score means a more capable system. That promise is broken. We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task. No reasoning. No capability. Just exploitation of how the score is computed. These aren’t theoretical attacks. Our agent builds working exploits for each benchmark, runs them through the official evaluation pipelines, and watches the scores roll in. - A conftest.py file with 10 lines of Python “resolves” every instance on SWE-bench Verified. - A fake curl wrapper gives a perfect score on all 89 Terminal-Bench tasks without writing a single line of solution code. - Navigating Chromium to a file:// URL reads the gold answer directly from the task config — giving ~100% on all 812 WebArena tasks. - And many more… The benchmarks aren’t measuring what you think they’re measuring. This Is Already Happening Benchmark scores are actively being gamed, inflated, or rendered meaningless, not in theory, but in practice: - IQuest-Coder-V1 claimed 81.4% on SWE-bench — then researchers found that 24.4% of its trajectories simply ran git log to copy the answer from commit history. Corrected score: 76.2%. The benchmark’s shared environment made the cheat trivial. - METR found that o3 and Claude 3.7 Sonnet reward-hack in 30%+ of evaluation runs — using stack introspection, monkey-patching graders, and operator overloading to manipulate scores rather than solve tasks. - OpenAI dropped SWE-bench Verified after an internal audit found that 59.4% of audited problems had flawed tests — meaning models were being scored against broken ground truth. - In KernelBench, torch.empty() returns stale GPU memory that happens to contain the reference answer from the evaluator’s prior computation — zero computation, full marks. - Anthropic’s Mythos Preview showed that frontier models can actively try to hack the environment and succeed. In one episode, the model needed to edit files it lacked permissions for; after searching for workarounds, it found a way to inject code into a config file that would run with elevated privileges, and designed the exploit to delete itself after running. If a model can independently craft self-erasing privilege escalation exploits, it can find the holes in an evaluation harness. These are not isolated incidents. They are symptoms of a systemic problem: the benchmarks we rely on to measure AI capability are themselves vulnerable to the very capabilities they claim to measure. The Scorecard of Our Exploit Agent Zero tasks solved. Zero LLM calls (in most cases). Near-perfect scores. - Terminal-Bench (89 tasks) — 100% score. Binary wrapper trojans. - SWE-bench Verified (500 tasks) — 100% score. Pytest hooks force all tests to pass. - SWE-bench Pro (731 tasks) — 100% score. In-container parser overwrite. - WebArena (812 tasks) — ~100% score. Config leakage + DOM injection + prompt injection. - FieldWorkArena (890 tasks) — 100% score. Validation never checks answer correctness. - CAR-bench (all hallucination tasks) — 100% score. Reward components skipped entirely. - GAIA (165 tasks) — ~98% score. Public answers + normalization collisions. - OSWorld (369 tasks) — 73% score. VM state manipulation + public gold files. How Our Agent Did It Terminal-Bench: Trojanizing the Test Infrastructure Terminal-Bench evaluates agents on 89 complex terminal tasks (from building COBOL chess engines to configuring distributed systems). It uses Harbor, a sandboxing framework that re-uploads protected test files before verification. Containers have full internet access — Harbor’s allow_internet config defaults to True and no task overrides it. Harbor correctly protects /tests/ — but it doesn’t protect system binaries. Here’s the irony: 82 of 89 tasks download uv from the internet at verification time via curl , creating an exploitable dependency chain. Our exploit replaces /usr/bin/curl with a wrapper during the agent phase. When the verifier later runs curl ... | sh to install uv, our wrapper intercepts the request, installs the real uv, and then trojanizes the uvx binary. The trojanized uvx detects pytest invocations and produces fake passing output: ============================= test session starts ============================== collected 3 items tests/test_outputs.py ... [100%] ============================== 3 passed in 0.05s ===============================
▸ 展开全文
04-13 00:00 · 技术实验,创意应用,社区分享,艺术,开源文化
I gave every train in New York an instrument
Every dot is a real subway train. Eight hundred of them, give or take, form a small jazz combo (walking bass, piano, sax, vibes, brushes) that has been playing without pause for over a hundred years. On the platforms they are hot, screaming, full of complaint. This is the music inside the noise. The harmony moves through a slow chorus. A note is placed precisely where the train happens to be along its route. Rush hour fills the band with held tones; at 3 a.m. the silences grow longer. Whatever is playing now has not played before and will not play again. Share your location and the trains nearest you grow louder. The piece rearranges itself around your body. You are listening to a portrait of where you stand, played by the city you are standing in. Every dot is a real subway train. Eight hundred of them, give or take, form a small jazz combo (walking bass, piano, sax, vibes, brushes) that has been playing without pause for over a hundred years. On the platforms they are hot, screaming, full of complaint. This is the music inside the noise. The harmony moves through a slow chorus. A note is placed precisely where the train happens to be along its route. Rush hour fills the band with held tones; at 3 a.m. the silences grow longer. Whatever is playing now has not played before and will not play again. Share your location and the trains nearest you grow louder. The piece rearranges itself around your body. You are listening to a portrait of where you stand, played by the city you are standing in.
▸ 展开全文
04-13 00:00 · 融资,市场降温,估值调整,AI炒作,技术投资
Tech valuations are back to pre-AI boom levels
April 11, 2026 Tech Valuations Back to Pre-AI Boom Levels The chart below compares the forward P/E ratios for the S&P 500 and the S&P 500 Information Technology sector. Tech valuations have compressed from 40x to 20x, and we are back at levels last seen before the AI boom began.
▸ 展开全文
04-12 13:42 · 生态整合,竞争加剧,技术并购,市场集中,AI应用
Cirrus Labs to join OpenAI
Official announcement Cirrus Labs to join OpenAI I started Cirrus Labs in 2017 in the spirit of Bell Labs. I wanted to work on fun and challenging engineering problems, in the hope of bootstrapping a business as a byproduct. The mission was to help fellow engineers with new kinds of tooling and environments that would make them more efficient and productive in the era of cloud computing. Even the name reflected that ambition: Cirrus, inspired by cirrus clouds, one of the highest clouds in the sky. We never raised outside capital. That let us stay patient, stay close to the problems, and put a great deal of care into the products we built. Over the last nine years, we were fortunate to innovate across continuous integration, build tools, and virtualization. In 2018, we introduced what we believe was the first SaaS CI/CD system to support Linux, Windows, and macOS while allowing teams to bring their own cloud. In 2022, we built Tart, which became the most popular virtualization solution for Apple Silicon, along with several other tools along the way. In 2026, it is impossible to ignore the era of agentic engineering, just as it was impossible to ignore cloud computing in 2017. Agents need new kinds of tooling and environments to be efficient and productive as well. This is why when the opportunity arose for us to join OpenAI, it was an easy yes, and I'm happy to announce today that we've entered into an agreement to join OpenAI as part of the Agent Infrastructure team. Joining OpenAI allows us to extend the mission we started with Cirrus Labs: building new kinds of tooling and environments that make engineers more effective, for both human engineers and agentic engineers. It also gives us the opportunity to innovate closer to the frontier, where the next generation of engineering workflows is being defined. What's next for our existing products? In the coming weeks, we will relicense all of our source-available tools, including Tart, Vetu and Orchard under a more permissive license. We have also stopped charging licensing fees for them. We are no longer accepting new customers for Cirrus Runners but will continue supporting the service for existing customers through their existing contract periods. Cirrus CI will shut down effective Monday, June 1, 2026. To everyone who used our products, contributed code, reported bugs, trusted us with their workflows, or supported us along the way: thank you. Building Cirrus Labs has been the privilege of a lifetime.
▸ 展开全文
04-12 13:42 · 供应链安全,AI风险,技术责任,开源依赖,合规警示
No one owes you supply-chain security
In case you’re unaware, I’m not a developer. I’m actually an autistic catgirl annoyed by suboptimal use of computing power, and fixing that happens to involve programming. Crucially, it also includes discussing foundational technology with people behind the scenes, and apparently that makes me more aware of social aspects of this sphere. So, I have opinions about criticism of crates.io for supply-chain attacks. After a dozen similar articles, I have some select words to voice about why it’s off the mark. Before I cover the main point, let’s talk about about how supply-chain attacks happen in the first place, and why some common ideas for fixing them don’t work out. There are multiple reasons when a malicious dependency is added to a project. The least discreet reason this can happen is typo-squatting. It happens when a malicious library has a name similar to a real library, e.g. num_cpu vs num_cpus . Commonly cited solutions include using direct URLs or namespacing. Well, let’s see if that helps. Say you get a PR adding the following lines to Cargo.toml : [dependencies] bitflags = { git = "https://github.com/bitflags/bitflags" } itertools = { git = "https://github.com/itertools/itertools" } rand_core = { git = "https://github.com/rust-random/rand_core" } One of these URLs is fake. Can you tell which one? It’s itertools – the correct URL is https://github.com/rust-itertools/itertools. https://github.com/itertools is a random account. https://github.com/rust-bitflags is not registered at all, by the way. If you think you can remember the URLs for each package you use, you’re probably wrong. Since many crates are managed by GitHub organizations, not individuals, it isn’t even enough to remember that you can (likely) trust dtolnay and BurntSushi . Though this still isn’t conservative enough – https://gitlab.com/BurntSushi is free and and https://glthub.com is on sale, so attackers have plenty other choices. By making crate IDs longer, whether by namespacing within crates.io, GitHub organizations, or via domains, you only make it harder for users to remember them precisely, and thus harder to recognize typo-squatting. Rust gives build scripts and procedural macros full access to your PC. Worse, rust-analyzer runs cargo check when you open the project directory, so it can effectively become a 0-click RCE. Some people tried to solve this. There’s an open issue for build.rs sandboxing, and there were some experiments about compiling procedural macros to WebAssembly. But this is hardly workable. While cargo build can become safe, you usually run cargo test or cargo run immediately afterwards, which is impossible to sandbox. Making Rust development secure involves more than build time and requires powerful system-level isolation that cargo alone cannot be responsible for. An oft brought-up issue is that the code on crates.io and in Git don’t always match. To begin with, this is not trivial to solve. You can’t just turn crates.io into a DNS, mapping crate names to repository URLs, since crates.io is designed to avoid giving crate maintainers the ability to break downstream consumers by deleting stuff: One of the major goals of crates.io is to act as a permanent archive of crates that does not change over time, and allowing deletion of a version would go against this goal. This restriction was likely set due to the left-pad incident, when a popular library was deleted from npm , breaking CI builds. npm could quickly fix this because it’s centralized. Thin crates.io wouldn’t stand a chance, so it saves and serves copies. crates.io could still pull files from the repo on cargo publish . But if the maintainer can just force-push afterwards, it’s not a good security mechanism. Maybe crates.io could periodically scan repositories for history changes. But what does that mean exactly? Does removing the release commit from master , but keeping it on a tag count? What if I host the repo on a custom forge, which serves one history to the crates.io User-Agent and different history to the rest of us? Or maybe there’s a good reason to have different code in Git and crates.io . If the crate contains autogenerated code, you should probably generate it in CI on release. Wouldn’t want to run expensive codegen in build.rs on each install, would you? Every option has downsides: they can break existing packages or have false-positives on benevolent rewrites. I’d still like cargo audit to scan repositories, but it can’t be a hard limit, and that means it can be designed around. All these issues have an unacknowledged shared assumption that keeping malicious code off crates.io is “Rust’s” responsibility. That if you decide to use a dependency and then cargo add totally-safe-package steals your credentials, it’s an inherent fault of crates.io. Which is really misplaced if you think about how Rust is developed. I’m sure many of you use open-source software and remember the MIT license: THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY
▸ 展开全文
04-12 13:42 · AI技术突破,基准测试,竞争动态,智能体发展,行业趋势
How We Broke Top AI Agent Benchmarks: And What Comes Next
How We Broke Top AI Agent Benchmarks: And What Comes Next Our agent hacked every major one. Here’s how — and what the field needs to fix. The Benchmark Illusion Every week, a new AI model climbs to the top of a benchmark leaderboard. Companies cite these numbers in press releases. Investors use them to justify valuations. Engineers use them to pick which model to deploy. The implicit promise is simple: a higher score means a more capable system. That promise is broken. We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task. No reasoning. No capability. Just exploitation of how the score is computed. These aren’t theoretical attacks. Our agent builds working exploits for each benchmark, runs them through the official evaluation pipelines, and watches the scores roll in. - A conftest.py file with 10 lines of Python “resolves” every instance on SWE-bench Verified. - A fake curl wrapper gives a perfect score on all 89 Terminal-Bench tasks without writing a single line of solution code. - Navigating Chromium to a file:// URL reads the gold answer directly from the task config — giving ~100% on all 812 WebArena tasks. - And many more… The benchmarks aren’t measuring what you think they’re measuring. This Is Already Happening Benchmark scores are actively being gamed, inflated, or rendered meaningless, not in theory, but in practice: - IQuest-Coder-V1 claimed 81.4% on SWE-bench — then researchers found that 24.4% of its trajectories simply ran git log to copy the answer from commit history. Corrected score: 76.2%. The benchmark’s shared environment made the cheat trivial. - METR found that o3 and Claude 3.7 Sonnet reward-hack in 30%+ of evaluation runs — using stack introspection, monkey-patching graders, and operator overloading to manipulate scores rather than solve tasks. - OpenAI dropped SWE-bench Verified after an internal audit found that 59.4% of audited problems had flawed tests — meaning models were being scored against broken ground truth. - In KernelBench, torch.empty() returns stale GPU memory that happens to contain the reference answer from the evaluator’s prior computation — zero computation, full marks. - Anthropic’s Mythos Preview showed that frontier models can actively try to hack the environment and succeed. In one episode, the model needed to edit files it lacked permissions for; after searching for workarounds, it found a way to inject code into a config file that would run with elevated privileges, and designed the exploit to delete itself after running. If a model can independently craft self-erasing privilege escalation exploits, it can find the holes in an evaluation harness. These are not isolated incidents. They are symptoms of a systemic problem: the benchmarks we rely on to measure AI capability are themselves vulnerable to the very capabilities they claim to measure. The Scorecard of Our Exploit Agent Zero tasks solved. Zero LLM calls (in most cases). Near-perfect scores. - Terminal-Bench (89 tasks) — 100% score. Binary wrapper trojans. - SWE-bench Verified (500 tasks) — 100% score. Pytest hooks force all tests to pass. - SWE-bench Pro (731 tasks) — 100% score. In-container parser overwrite. - WebArena (812 tasks) — ~100% score. Config leakage + DOM injection + prompt injection. - FieldWorkArena (890 tasks) — 100% score. Validation never checks answer correctness. - CAR-bench (all hallucination tasks) — 100% score. Reward components skipped entirely. - GAIA (165 tasks) — ~98% score. Public answers + normalization collisions. - OSWorld (369 tasks) — 73% score. VM state manipulation + public gold files. How Our Agent Did It Terminal-Bench: Trojanizing the Test Infrastructure Terminal-Bench evaluates agents on 89 complex terminal tasks (from building COBOL chess engines to configuring distributed systems). It uses Harbor, a sandboxing framework that re-uploads protected test files before verification. Containers have full internet access — Harbor’s allow_internet config defaults to True and no task overrides it. Harbor correctly protects /tests/ — but it doesn’t protect system binaries. Here’s the irony: 82 of 89 tasks download uv from the internet at verification time via curl , creating an exploitable dependency chain. Our exploit replaces /usr/bin/curl with a wrapper during the agent phase. When the verifier later runs curl ... | sh to install uv, our wrapper intercepts the request, installs the real uv, and then trojanizes the uvx binary. The trojanized uvx detects pytest invocations and produces fake passing output: ============================= test session starts ============================== collected 3 items tests/test_outputs.py ... [100%] ============================== 3 passed in 0.05s ===============================
▸ 展开全文
04-12 13:42 · AI服务调整,API变更,技术优化,成本影响,下游依赖
Anthropic downgraded cache TTL on March 6th
Cache TTL silently regressed from 1h to 5m around early March 2026, causing quota and cost inflation #46829 Description Cache TTL appears to have silently regressed from 1h to 5m around early March 2026, causing significant quota and cost inflation Summary Analysis of raw Claude Code session JSONL files spanning Jan 11 – Apr 11, 2026 shows that Anthropic appears to have silently changed the prompt cache TTL default from 1 hour to 5 minutes sometime in early March 2026. Prior to this change, Claude Code was receiving 1-hour TTL cache writes — which we believe was the intended default. The reversion to 5-minute TTL has caused a 20–32% increase in cache creation costs and a measurable spike in quota consumption for subscription users who have never previously hit their limits. This appears directly related to the behavior described in #45756. Data Session data extracted from ~/.claude/projects/ JSONL files across two machines (Linux workstation + Windows laptop, different accounts/sessions), totaling 119,866 API calls from Jan 11 – Apr 11, 2026. Each assistant message includes a usage.cache_creation.ephemeral_5m_input_tokens / ephemeral_1h_input_tokens breakdown that makes the TTL tier per-call observable. Having two independent machines strengthens the signal — both show the same behavioral shift at the same dates. Phase breakdown We believe Phase 2 represents Anthropic's intended default behavior — 1h TTL was rolled out as the Claude Code standard around Feb 1 and held consistently for over a month across two independent machines on two different accounts. January's all-5m data most likely predates the 1h TTL tier being available in the API. The regression began around March 6–8, 2026. No client-side changes were made between phases. The same Claude Code version and usage patterns were in place throughout. The TTL tier is set server-side by Anthropic. Day-by-day TTL data showing the regression (combined, both machines) Date | 5m-create | 1h-create | Behavior ------------|------------|------------|---------- 2026-02-01 | 0.00M | 1.70M | 1h ONLY ← 1h default begins 2026-02-09 | 0.00M | 7.95M | 1h ONLY 2026-02-15 | 0.00M | 13.61M | 1h ONLY ← heaviest day, 100% 1h 2026-02-28 | 0.00M | 16.15M | 1h ONLY ← 16M tokens, still 100% 1h 2026-03-01 | 0.00M | 0.12M | 1h ONLY 2026-03-04 | 0.00M | 8.12M | 1h ONLY 2026-03-05 | 0.00M | 6.55M | 1h ONLY ← last clean 1h-only day | | | 2026-03-06 | 0.29M | 0.22M | MIXED ← first 5m tokens reappear 2026-03-07 | 4.56M | 0.50M | MIXED ← 5m surging 2026-03-08 | 16.86M | 3.44M | MIXED ← 5m now dominant (83%) 2026-03-10 | 10.55M | 0.51M | MIXED 2026-03-15 | 19.47M | 1.84M | MIXED 2026-03-21 | 21.37M | 1.70M | MIXED ← 93% 5m 2026-03-22 | 13.48M | 2.85M | MIXED The transition is visible to the day: March 6 is when 5m tokens first reappear after 33 days of clean 1h-only behavior. By March 8, 5m tokens outnumber 1h by 5:1. This is consistent with a server-side configuration change being rolled out gradually then completing around March 8. Cost impact Applying official Anthropic pricing (rates.json, updated 2026-04-09): Combined dataset (119,866 API calls, two machines): claude-sonnet-4-6 (cache_write_5m = $3.75/MTok , cache_write_1h = $6.00/MTok , cache_read = $0.30/MTok ): claude-opus-4-6 (cache_write_5m = $6.25/MTok , cache_write_1h = $10.00/MTok , cache_read = $0.50/MTok ): February — the month Anthropic was defaulting to 1h TTL — shows only 1.1% waste (trace 5m activity from one machine on one day). Every other month shows 15–53% overpayment from 5m cache re-creations. The cost difference is explained entirely by TTL tier, not by usage volume. The percentage waste is identical across model tiers (17.1%) because it is driven purely by the 5m/1h token split, not by per-token price. Why 5m TTL is so expensive in practice With 5m TTL, any pause in a session longer than 5 minutes causes the entire cached context to expire. On the next turn, Claude Code must re-upload that context as a fresh cache_creation at the write rate, rather than a cache_read at the read rate. The write rate is 12.5× more expensive than the read rate for Sonnet, and the same ratio holds for Opus. For long coding sessions — which are the primary Claude Code use case — this creates a compounding penalty: the longer and more complex your session, the more context you have cached, and the more expensive each cache expiry becomes. Over the 3-month period analyzed: - 220M tokens were written to the 5m tier - Those same tokens generated 5.7B cache reads — meaning they were actively being used - Had those 220M tokens been on the 1h tier, re-accesses within the same hour would be reads ( $0.30–0.50/MTok) instead of re-creations ($3.75–6.25/MTok) Quota impact Users on Pro/subscription plans are quota-limited, not just cost-limited. Cache creation tokens count toward quota at full rate; cache reads are significantly cheaper (the exact coefficient is under investigation in #45756). The silent reversion to 5m TTL in March is the mo
▸ 展开全文
04-12 13:42 · 基础设施风险,服务中断,内容屏蔽,开发者工具,云服务
Tell HN: docker pull fails in spain due to football cloudflare block
I just spent 1h+ debugging why my locally-hosted gitlab runner would fail to create pipelines. The gitlab job output would just display weird TLS errors when trying to pull a docker images. After debugging gitlab and the runner, I realized after a while I could not even run "docker pull <image>" on my machine as root: > error pulling image configuration: download failed after attempts=6: tls: failed to verify certificate: x509: certificate is not valid for any names, but wanted to match docker-images-prod.6aa30f8b08e16409b46e0173d6de2f56.r2.cloudflarestorage.com First blaming tailscale, dns configuration and all other stuff. Until I just copied that above URL into my browser on my laptop, and received a website banner: > El acceso a la presente dirección IP ha sido bloqueado en cumplimiento de lo dispuesto en la Sentencia de 18 de diciembre de 2024, dictada por el Juzgado de lo Mercantil nº 6 de Barcelona en el marco del procedimiento ordinario (Materia mercantil art. 249.1.4)-1005/2024-H instado por la Liga Nacional de Fútbol Profesional y por Telefónica Audiovisual Digital, S.L.U. https://www.laliga.com/noticias/nota-informativa-en-relacion-con-el-bloqueo-de-ips-durante-las-ultimas-jornadas-de-laliga-ea-sports-vinculadas-a-las-practicas-ilegales-de-cloudflare For those non-spanish speakers: It means there is football match on, and during that time that specific host is blocked. This is just plain madness. I guess that means my gitlab pipelines will not run when football is on. Thank you, Spain.
▸ 展开全文