03-29 00:01 · AI,技术,HackerNews,AI
CERN uses ultra-compact AI models on FPGAs for real-time LHC data filtering
[ GENEVA, SWITZERLAND — March 28, 2026 ] — CERN is using extremely small, custom artificial intelligence models physically burned into silicon chips to perform real-time filtering of the enormous data generated by the Large Hadron Collider (LHC).
OVERVIEW
The Large Hadron Collider (LHC) generates an extraordinary volume of raw data — approximately 40,000 exabytes per year, equivalent to roughly one quarter of the entire current internet. During peak operation, the data stream can reach hundreds of terabytes per second, far exceeding the capacity of any feasible storage or conventional computing system.
Because it is physically impossible to store or process the full dataset, CERN must make split-second decisions at the detector level: which collision events contain potentially groundbreaking scientific value, and which should be discarded forever. This real-time selection process is one of the most demanding computational challenges in modern science.
To meet these extreme requirements, CERN has deliberately moved away from conventional GPU or TPU-based artificial intelligence architectures. Instead, the laboratory develops highly optimized, ultra-compact AI models that are compiled and physically implemented directly into custom silicon — primarily field-programmable gate arrays (FPGAs) and application-specific integrated circuits (ASICs). These hardware-embedded models enable ultra-low-latency inference at the very edge of the detector system, where decisions must be made in microseconds or even nanoseconds.
THE DATA CHALLENGE
Inside the 27-kilometre ring of the Large Hadron Collider, proton bunches travel at velocities approaching the speed of light and cross paths roughly every 25 nanoseconds. Although billions of protons pass through one another during each crossing, actual hard collisions between protons remain relatively rare events.
When a collision does occur, the detectors surrounding the interaction point capture several megabytes of raw data from the resulting particle shower. This creates an overwhelming data stream: the LHC can generate up to hundreds of terabytes per second at peak luminosity. Storing or processing the full volume is physically impossible with current technology.
As a result, only about 0.02 % of all collision events are ultimately retained for further analysis. The first and most critical filtering stage, known as the Level-1 Trigger, is responsible for making these split-second decisions. It consists of approximately 1,000 field-programmable gate arrays (FPGAs) that evaluate incoming data in less than 50 nanoseconds. A highly specialized algorithm called AXOL1TL runs directly on these chips, analysing detector signals in real time and determining which events are scientifically promising enough to be preserved. All other data is discarded immediately and permanently.
AI APPROACH AND TECHNICAL STACK
CERN’s artificial intelligence models are deliberately designed to be extremely small and highly optimised for the unique constraints of the LHC environment. Unlike the large-scale language models and general-purpose AI systems commonly used in industry, these models are tailored specifically for ultra-low-latency, real-time inference at the detector level, where decisions must be made in nanoseconds.
The models are compiled using the open-source tool **HLS4ML**, which translates machine-learning models written in frameworks such as PyTorch or TensorFlow into synthesizable C++ code. This code can then be deployed directly onto field-programmable gate arrays (FPGAs), systems-on-chip (SoCs), or custom application-specific integrated circuits (ASICs). The resulting hardware implementations achieve the extreme speed required while consuming significantly less power and silicon area than conventional GPU- or TPU-based solutions.
A distinctive feature of CERN’s approach is that a substantial portion of the available chip resources is not allocated to the neural network layers themselves. Instead, these resources are used to implement extensive precomputed lookup tables. These tables store the results of common input patterns in advance, allowing the hardware to deliver near-instantaneous outputs for the vast majority of typical detector signals without performing full floating-point calculations. This hardware-first design philosophy is what enables the system to operate at the required nanosecond-scale latency.
The second filtering stage, known as the High-Level Trigger, runs on a large surface-level computing farm consisting of 25,600 CPUs and 400 GPUs. Even after the aggressive Level-1 Trigger has reduced the data volume, this farm must still process terabytes of data per second before further reducing it to approximately one petabyte of scientifically valuable data per day.
FUTURE PLANS
The current Large Hadron Collider is scheduled for a major upgrade known as the High-Luminosity LHC (HL-LHC), which is expected to begin operations in 2031. This upgrade will dramatically increase t
▸ 展开全文
03-29 00:01 · AI,技术,HackerNews,AI
InpharmD (YC W21) Is Hiring – Senior Ruby on Rails Developer
| InpharmD Jobs
Senior Ruby on Rails Engineer
Senior Ruby on Rails Engineer — InpharmD
At InpharmD, we help healthcare providers make better clinical decisions by giving them evidence-backed answers to their questions.
Founded: 2018
Funding: Seed ($8.05M)
We’ve grown revenue 750% while staying capital efficient. Small team. High ownership. No politics. Just people who want to build meaningful healthcare products fast.
Why InpharmD
- Capital efficient growth — we don’t burn, we build
- Small, high-performing team — no bloated org layers
- Ownership mindset — builders, not passengers
- Fast iteration — we ship quickly and improve constantly
- No drama culture — we meet once a week, not every day
Market timing is in our favor — and we’re moving faster than incumbents.
Role: Senior Ruby on Rails Engineer (10+ years)
We’re looking for a strong backend engineer to help us scale our core platform and APIs that power clinical decision-making.
This is a hands-on role. You’ll design systems, write production-grade code, and own key parts of our backend.
What You’ll Do
- Build and scale Ruby on Rails (Rails 8+) APIs
- Develop systems using Ruby 3+ with modern patterns
- Design clean, scalable database architectures (Postgres, large datasets)
- Build and manage background processing using Sidekiq or Solid Queue
- Work on high-performance systems handling clinical data and workflows
- Integrate healthcare pricing systems (340B, WAC)
- Improve system reliability, performance, and maintainability
- Collaborate closely with AI/ML teams powering our core product
What We’re Looking For
- 10+ years of experience with Ruby on Rails (production systems)
- Strong experience with Rails 8+ and Ruby 3+
- Hands-on experience with Sidekiq or Solid Queue (job orchestration, scaling workers)
- Strong experience designing APIs and backend architectures
- Deep understanding of databases, data modeling, and performance optimization
- Experience working with large datasets and distributed systems
- Familiarity with healthcare systems (340B / WAC pricing) is a strong plus
- You care about clean code, speed, and real-world impact
- You take ownership and consistently deliver high-quality work
Tech Stack (Relevant)
- Ruby on Rails (Rails 8+)
- Ruby 3+
- PostgreSQL
- Sidekiq / Solid Queue
- AWS (S3, EC2, etc.)
- API-first architecture
Logistics
- Location: Atlanta Tech Village (preferred) or Remote
- Compensation: $130K base + InpharmD stock options
- Hours: Full-time
- Contact: Tulasee Rao Chintha
Equal Opportunity
We value diverse perspectives and strong builders. If you’re great at what you do, we want to talk.
Interested?
Learn more about us on our InpharmD blog.
If you’re excited about building the future of AI in healthcare, send an email to founders@inpharmd.com — your message goes directly to the founders. Feel free to include why you’re interested and what kind of opportunity you’re looking for.
Share
▸ 展开全文
03-29 00:01 · AI,技术,HackerNews,AI
The first 40 months of the AI era
The first 40 months of the AI era
Here are my accumulated thoughts and ideas about AI since the launch of ChatGPT in November 2022.
40 months of ChatGPT
OpenAI first launched ChatGPT at the end of November 2022, nearly 40 months ago. I remember trying it at the time and being really impressed, just like everyone else. At first I just talked to it, and I remember being absolutely amazed. I remember what it was like to talk to older chat bots, more primitive programs like Cleverbot. ChatGPT was better, much better. So much so that it was immediately obvious that this was not going to be just another toy for internet nerds, the rest of the world was going to notice.
I experimented with asking it to write, to create content. I first asked it to create poems, then Dungeons and Dragons backgrounds and even a whole fantasy world including important characters, kingdoms, and lore. It was very impressive, just for the fact that the output was coherent but its ‘style’ was very boring and overtly inoffensive, which was (and still is) a clear limitation of the technology.
I was listening to the Linus Tech Tips WAN Show immediately following the launch and hearing Luke mention that ChatGPT was being prompted to produce fully functional programs. This made me very curious and I wanted to try prompting ChatGPT to test its capabilities. I first asked for simple hello world programs, which it produced perfectly. I was very impressed and continued to play with it, and it didn’t take long for me to realise that this bot could produce genuinely useful code snippets for common and well understood use cases. For simple things it was able to replace my typical research loop, I no longer needed to search on StackOveflow or other discussion forums for these kinds of solutions.
I remember the first time I vibe-coded a small project. It was an app that generated placeholder cards for my MTG collection. I prompted the bot (now Claude, not ChatGPT) to create an app that would fetch for the card metadata from an API, generate a qrcode, and correctly layout this information into a printable page of cards. The first output was very impressive, and mostly worked. I attempted to make adjustments with additional prompting but was unable to make meaningful progress. I gave up on using the bot, and finished the project on my own. Through each iteration of the project I replaced more of what the bot wrote with my own code. The final result hardly used any AI generated code at all. I debated with myself whether I had actually saved any time or effort, compared with having just done everything myself from the very beginning. Coding AI has come a long way since then, but even now I’m constantly asking myself just how much is this actually useful for coding?
Claude Code, my Review
Two months ago I purchased a Claude Pro subscription for the first time. Claude had been my free chat bot of choice for over a year and I had become increasingly curious about Claude Code. My initial impression was intensely positive. I immediately installed Claude Code on my workstation and began experimenting. I remember the initial excitement from being able to speak naturally to my computer. So long as I was careful to clarify my intent, I was now able to tell my computer what I wanted it to do and it would consistently do what I asked. It felt then (and now) like a brand new form of input and control of my computer along side my keyboard, mouse, and even command line terminal. I have a lot of doubts about using AI, but not for this use case. This is unambiguously good, useful, and just amazing. I would be very pleased to see this technology become fully commoditized. I would love to have a local LLM loaded on my GPU, or on a separate desktop appliance, that can do just this, and do so this well.
And of course I also tried vibe coding using Claude Code. The results were again very impressive. For the small projects that I attempted, I was able to get a good one-shot result (just like I did previously). This time the iterative prompting felt much more productive. The Claude Code interface eliminates the friction of copy/pasting from a chat interface, the bot can just make the edit itself. I was impressed with model’s ability to maintain coherency and context. I was amazed when it created solutions or found bugs that I failed to see. But even on seemingly simple projects it felt like I was struggling to keep it from eventually losing the plot.
I also tried to use Claude Code to help me start a business. Since I lost my job as an IT Technician last year I had considered building my own small IT services company. I tried to use Claude as a combination of executive assistant and mentor. I asked it to create a detailed pre-launch plan that I could follow, and then to track my progress. In hindsight it all seems painfully obvious and simple, but I must admit that the mere process of creating the plan was very inspiring and created a lot of confidence in me. I did manage to l
▸ 展开全文
03-28 00:01 · AI,技术,HackerNews,AI
Everything old is new again: memory optimization
At this point in history, AI sociopaths have purchased all the world's RAM in order to run their copyright infringement factories at full blast. Thus the amount of memory in consumer computers and phones seems to be going down. After decades of not having to care about memory usage, reducing it has very much become a thing.
Relevant questions to this state of things include a) is it really worth it and b) what sort of improvements are even possible. The answers to these depend on the task and data set at hand. Let's examine one such case. It might be a bit contrived, unrepresentative and unfair, but on the other hand it's the one I already had available.
Suppose you have to write script that opens a text file, parses it as UTF-8, splits it into words according to white space, counts the number of time each word appears and prints the words and counts in decreasing order (most common first).
The Python baseline
This sounds like a job for Python. Indeed, an implementation takes fewer than 30 lines of code. Its memory consumption on a small text file [update: repo's readme, which is 1.3k] looks like this.
Peak memory consumption is 1.3 MB. At this point you might want to stop reading and make a guess on how much memory a native code version of the same functionality would use.
The native version
A fully native C++ version using Pystd requires 60 lines of code to implement the same thing. If you ignore the boilerplate, the core functionality fits in 20 lines. The steps needed are straightforward:
- Mmap the input file to memory.
- Validate that it is utf-8
- Convert raw data into a utf-8 view
- Split the view into words lazily
- Compute the result into a hash table whose keys are string views, not strings
The main advantage of this is that there are no string objects. The only dynamic memory allocations are for the hash table and the final vector used for sorting and printing. All text operations use string views , which are basically just a pointer + size.
In code this looks like the following:Its memory usage looks like this.
Peak consumption is ~100 kB in this implementation. It uses only 7.7% of the amount of memory required by the Python version.
Isn't this a bit unfair towards Python?
In a way it is. The Python runtime has a hefty startup cost but in return you get a lot of functionality for free. But if you don't need said functionality, things start looking very different.
But we can make this comparison even more unfair towards Python. If you look at the memory consumption graph you'll quite easily see that 70 kB is used by the C++ runtime. It reserves a bunch of memory up front so that it can do stack unwinding and exception handling even when the process is out of memory. It should be possible to build this code without exception support in which case the total memory usage would be a mere 21 kB. Such version would yield a 98.4% reduction in memory usage.
▸ 展开全文
03-28 00:01 · AI,技术,HackerNews,AI
Why are executives enamored with AI, but ICs aren't?
I think there’s pretty clearly a divide in AI perception between executives and individual contributors (ICs). Executives seem to love it and evangelize it (going so far as to creating mandates at their companies for AI usage). But ICs are typically much more skeptical of its usage. You can see the divide show up everywhere from Hacker News comment threads to internal Slack debates about adopting coding agents.
Here’s my current posit for why there’s such a big divide: executives have always had to deal with non-determinism and focus on nondeterministic system design, while individual contributors are evaluated by their execution on deterministic tasks.
Managing non-deterministic systems
Executives have always had to deal with non-determinism. That’s par for the course:
- People being out sick or taking time off unexpectedly
- Someone not finishing an important project and not talking about it until far too late in the process
- People reacting to an announcement in an unexpected way
- A feature being built in a way that doesn’t make sense with respect to the rest of the product, but does technically achieve objectives.
More generally, if you’ve ever taken a Chaos Theory class in math, you’ll know that nonlinear, chaotic systems emerge when individual agents in a system are all acting with different inputs, utility functions, etc. Systems become slightly easier to manage if you’re able to make those utility functions consistent (you’re able to get a grasp on system dynamics).
A manager’s job is to create a model of the world and align everyone’s utility functions, knowing that there’s a large amount of non-determinism in complex systems. So it makes sense that as a manager, you’re ok with a decent amount of this.
AI is something that is non-deterministic but has a lot of characteristics of a well behaved chaotic system (specifically a system where you can understand the general behavior of the system, even if you cannot predict the specific outcomes at any point in time).
For example:
- LLMs generally continue their work and provide an output regardless of time of day, how difficult the task is, how much information is available
- LLM’s deficiencies have well defined failure modes (e.g. hallucinations, lack of ability to operate outside of their context, and especially poor outcomes when not given enough context)
- The types of tasks that an LLM can accomplish are relatively well known, and the capability envelope is getting mapped out quickly. This is different than humans, where each person has a different set of strengths and weaknesses and where you need to uncover these over time.
Many of these properties are more deterministic than large human systems, which makes AI incredibly attractive for an executive who is already used to this and likely has put a large amount of effort into adding determinism into their systems already (e.g. by adding processes and structure in the form of levels and ladders, standard operating procedures, etc.).
ICs live in a more deterministic world
ICs are generally much more focused on particular problems that have specific inputs and outcomes. Correctness is easier to determine, and how good you are at your job can largely be described by quality and speed, where the weights on those two depend on which organization you’re in. This changes as you move up the ladder (a staff engineer is expected to tackle large, ambiguous business problems), but for most ICs, the world is relatively well defined.
ICs deal with plenty of non-determinism in practice (unclear requirements, flaky systems, shifting priorities), but the way they’re evaluated pushes in the other direction. An IC’s value often comes from being reliably precise (e.g. writing correct code, getting the analysis right, producing a design that holds up under scrutiny). The more deterministic your output, the better you are at your job.
AI introduces non-determinism into exactly this space, and from an IC’s perspective, there are good reasons to be skeptical:
- It’s not as good as they are at their job. A highly trained human focused on a specific task will often beat an LLM, especially if that task is long running, requires connecting multiple systems, or demands precise domain intuition. If you’re an expert and you’re handed a tool that does a mediocre version of your work, the overhead of fixing its mistakes can genuinely cost more than doing it yourself.
- It changes what their job is. You go from doing the work yourself to managing something that does the work. The skills that got you hired (deep focus, precision, domain knowledge) aren’t necessarily the skills that make you good at that. That’s a disorienting shift.
- It’s tied to self worth. Work accounts for the majority of a person’s waking hours. When executives talk about AI making everyone more productive, ICs can hear that as the things you’ve spent years getting good at are about to matter less. Whether or not that’s what’s actually being said, it’s a reaso
▸ 展开全文
03-28 00:01 · AI,技术,HackerNews,AI
21,864 Yugoslavian .yu domains
TLDR; get a list of 21,864 domains from the former Yugoslavia’s “.yu” top level domain: download the .CSV
In 2010 the entire domain space of Yugoslavia (.yu) was taken off the internet. After all, the country didn’t exist anymore.
I heard about this from an interview with Kaloyan Kolev on Agnes Bytes’ “Archiving the Web”. Kaloyan had several interesting insights:
- We’ve baked the concept of “countries” into the Internet domain system. And that makes domain names tied to real-world territorial conflicts, countries splitting apart and countries going underwater (like Tuvalu)
- .yu is an early example of something that will happen more and more often.
- It is unfortunate that we didn’t preserve the .yu domain space like a nostalgic Internet memorial to the country. Instead, the sites became unmoored and unreachable.
Kaloyan referred to the research paper “What does the Web remember of its deleted past? An archival reconstruction of the former Yugoslav top-level domain” by Anat Ben-David. In that paper, Ms. Ben-David reconstructed a network graph of .yu domains from the Internet Archive’s Wayback Machine.
Ben-David’s paper used the below sources as seed lists:
By crawling links from these pages to other .yu URLs, she eventually found 17,460 unique websites in the .yu domain.
The adventure begins
Dear Reader: after hearing all this, I bellowed out a mighty
Akshuallyyyyyyyy!
I dropped my bag of mini M&Ms onto the house robe I was wearing. Unshaven and red eyed, I yelled up from the basement: “MOM, fire up the router! I’m going on an Internet Adventure!!!!”.
You see, I figured I was good enough to discover all the archived domains under the .yu TLD. After all, last time I had an akshually moment, good things happened!
At first I tried doing a wildcard search for all *.yu domains at the Wayback Machine. That didn’t work.
The CDX API – a dead(ish) end
Then, I discovered that the Wayback Machine has a “CDX Server” API that can tell us if a page is archived or not.
Below is an example of a CDX query that grabs all the unique file paths at the domain “jacobfilipp.com”, filters them down only to HTML files that were successfully fetched (status starts with a 2), and shows you only the first 10 that are archived at the Wayback Machine.
https://web.archive.org/cdx/search/cdx?url=jacobfilipp.com&matchType=host&collapse=urlkey&filter=mimetype:text/html&filter=statuscode:^2&limit=10
Unfortunately you can’t easily fetch all archived URLs under the TLD “.yu”.
However, if you try, you get a message that says “Forbidden: This type of CDX query requires authorization.” Which tells me that you could do this if you politely ask the staff at the Internet Archive.
What does work is fetching all the files under the Yugoslavian subdomains like *.co.yu and *.org.yu and *.ac.yu.
Here is an example of fetching all the URLs under *.co.yu:
You’d need to paginate through all the results, and it’s slow.
WWW.YU to the rescue
While tinkering with the CDX API, I landed by mistake on the site “www.yu”. I believe this site was run by Yugoslavian ISP “Memodata”.
What’s special about it, is it has a list of just about every registered .yu domain:
I went ahead and used my newfound CDX skills to download a list of all indexed “domain listing” pages. Not all of them are in the Wayback Machine: most listings stop at “page 20” of each letter.
Then, I downloaded all those pages locally using wget (use the id_
URL trick to get a page with un-altered URLs), and extracted all the .yu domain names from the links inside.
Finally, looping through each domain name, I used the CDX endpoint to check whether the domain is in the Archive or not.
End result:
21,864 domains with 13,292 of them having an archived copy in the Wayback Machine.
Download the entire list as .CSV below:
An exercise for the reader
While writing this post, I realized that the parent of www.yu – memodata.net – also has a list of domains. Theoretically it is the same list of domains. But, practically, the Wayback Machine might have indexed alphabetical listing pages that it didn’t index for www.yu. You’d have to grab all the listings pages using the CDX API and extract the domains.
If you really need more .yu domains, Nikola Smolenski and Anat Ben-David are easy to find online. You should ask them nicely – I bet they have their lists saved somewhere.
▸ 展开全文
03-28 00:01 · AI,技术,HackerNews,AI
DOJ confirms FBI Director Kash Patel's personal email was hacked
Iran-linked hackers successfully broke into FBI Director Kash Patel’s personal email, the Department of Justice confirmed to Reuters on Friday.
Reuters could not authenticate the leaked emails themselves but noted that the Gmail address matched an email account “linked to Patel in previous data breaches preserved by the dark web intelligence firm District 4 Labs.” The DOJ suggested the emails appeared to be authentic.
On their website, the Handala Hack Team boasted that Patel “will now find his name among the list of successfully hacked victims.” The hacker group taunted Patel by sharing photos of him sniffing cigars and holding up a jug of rum, along with other documents that Reuters reported were from 2010 to 2019.
“Soon you will realize that the FBI’s security was nothing more than a joke,” the group posted, as documented in screenshots from the website shared widely on X.
The hack came after the DOJ disrupted some of the hacker group’s websites earlier this month. In a press release, Patel threatened to “hunt” down the group, which Reuters reported “calls itself a group of pro-Palestinian vigilante hackers.” After detailing four attacks this month that the group had taken credit for, Patel offered rewards of up to $10 million for information on its members.
“Iran thought they could hide behind fake websites and keyboard threats to terrorize Americans and silence dissidents,” Patel said. “We took down four of their operation’s pillars and we’re not done. This FBI will hunt down every actor behind these cowardly death threats and cyberattacks and will bring the full force of American law enforcement down on them.”
▸ 展开全文
03-27 05:47 · AI,技术,HackerNews,AI
Chroma Context-1: Training a Self-Editing Search Agent
Using search systems in conjunction with a large language model (LLM) is a common paradigm for enabling language models to access data beyond their training corpus. This approach, broadly known as retrieval-augmented-generation (RAG), has traditionally relied on single-stage retrieval pipelines composed of vector search, lexical search, or regular expression matching, optionally followed by a learned reranker. While effective for straightforward lookup queries, these pipelines are fundamentally limited: they assume that the information needed to answer a question can be retrieved in a single pass.
In practice, many real-world queries are not satisfiable in a single-stage. Answering a question often requires a chain of intermediate searches in which the output of one search informs the next, a process known as a multi-hop retrieval.
To solve this, leveraging LLMs for multi-turn agentic search has become a viable approach to answering multi-hop retrieval queries. Rather than issuing a single query, an LLM agent iteratively decomposes a high-level question into subqueries, retrieves evidence, and refines its search strategy across multiple turns. Concurrently, it has been shown that smaller-parameter language models, trained on moderate-scale corpora, can serve as effective search agents with performance comparable to substantially larger models. Running frontier-scale models for multi-turn search incurs high cost and latency, which motivates offloading this task to a smaller, purpose-trained model.
A key factor driving the cost and latency of agentic search is the growth of the context window. As the agent gathers information over multiple turns, its context window fills rapidly with retrieved documents, many of which may be tangential or redundant. This bloated context not only increases computational cost but can also degrade downstream performance due to increasing the presence of distracting information. One promising direction to address this is self-editing context, in which the agent actively decides which retrieved information to retain and which to discard, allowing it to continue long-horizon search tasks more efficiently and more accurately within a bounded context window.
Building on these insights, we trained Chroma Context-1, a 20B parameter agentic search model on over eight thousand synthetically generated tasks. Context-1 achieves retrieval performance comparable to frontier LLMs at a fraction of the cost and up to 10x the inference speed. Context-1 operates as a retrieval subagent: rather than answering questions directly, it returns a ranked set of supporting documents to a downstream answering model, cleanly separating search from generation. The model is trained to decompose a high-level query into subqueries and iteratively search a corpus across multiple turns. As the agent's context window fills, it selectively discards irrelevant results to free capacity and reduce noise for further exploration.
In this work we present our synthetic data generation pipeline, agent harness, and training methodology alongside a comprehensive evaluation of Context-1 across a range of retrieval benchmarks. Our results demonstrate that a purpose-trained 20B model can reach the Pareto frontier of retrieval performance with respect to cost and latency, matching or exceeding frontier models that are orders of magnitude larger at a fraction of the compute.
We present the following:
- A staged training curriculum that first optimizes for recall before shifting toward precision, training the agent to progressively narrow from broad retrieval to selective retention. We release the weights of this model to the public under a permissive Apache 2.0 license.
- A context management strategy in which the agent selectively edits its own context during search, discarding irrelevant passages to free context capacity for further exploration and to reduce the effects of context rot.
- A scalable synthetic task generation pipeline that uses a human-aligned LLM judge to minimize the need for human annotation while maintaining task quality. We release the full codebase for this pipeline to support reproducibility and further research.
The limitations of single-shot retrieval have driven substantial exploration into agentic search systems, in which reasoning is interleaved with retrieval to resolve queries that require satisfying multiple constraints jointly or following a chain of dependent clues across documents. These systems vary in their termination strategy: some run for a fixed number of turns, while others terminate dynamically based on a learned sufficiency signal. By shifting control of the retrieval strategy to the model itself, these systems can reformulate queries based on intermediate results, decide when to explore versus exploit, and terminate search based on a confidence assessment. These systems model search as a sequential reasoning task, in which the right next query depends on what has been found so far. Be
▸ 展开全文
03-27 05:47 · AI,技术,HackerNews,AI
Agent-to-agent pair programming
What if you could let Claude and Codex work together as pair programmers, talking to each other directly? One of them as the main worker and the other as a reviewer.
It is amusing how the best agentic workflows often look a lot like human collaboration. Researchers at Cursor discovered this in their work on long-running coding agents. That work led them to create a multi-agent workflow with a main orchestrator assigning tasks to workers. This is similar to how most human teams operate. Claude Code “Agent teams” and Codex “Multi-agent” features work similarly, with subagents reporting back to the main agent. And in the future, subagents could interact with each other, like humans do.
I wanted to pursue the idea of mimicking human collaboration with multiple agent harnesses and another workflow used by programmers: pair programming. While building a code review agent using Claude and Codex side-by-side, I found something interesting: they gave different feedback -- but even when they gave the same feedback, it wasn’t annoying: it was in fact a very strong signal. Our team addresses 100% of the feedback when both reviewers agree. Code reviews are great because they happen on a multiplayer app where humans and agents collaborate, but they are slowing down the feedback loop and can become noisy.
That’s why I built loop
: a dead-simple CLI that launches claude
and codex
side-by-side in tmux, with a bridge that lets them talk to each other. It makes this feedback loop faster and more natural, while preserving context across iterations. It’s interesting because it enables the agents to be more proactive, since the interaction between them is more natural (and I expect that to only get better as the models get better too). Because loop
runs the interactive TUIs, you can stay in the loop, steer, answer questions, and follow up if needed.
The future of agentic workflows may look less like magic automation and more like familiar teamwork. And I’m sure that there are some great observations to apply to this pair programming workflow. Some open questions around how to make the human handoff and PR review easier:
- Should we split the work across multiple PRs?
- Should we share the PLAN.md in git or in the PR description?
- Should we share a screenshot or video recording as a proof of work?
Letting the agents loop can result in more changes than expected, which are usually welcome -- but unfortunately it makes the human review harder.
A lot of people are using multiple agent harnesses for a variety of reasons: to avoid vendor lock-in, to use and contribute to an open-source project, to max out their subscriptions, or to get different perspectives, strengths, and results. Multi-agent harness apps should probably treat agent-to-agent communication as a first-class feature. I’d love to see them adopt this approach.
Try it out: https://github.com/axeldelafosse/loop
Thanks to Léna Deloizy Delafosse, Will Horn, Tian Wang and Ferruccio Balestreri for reading drafts of this.
▸ 展开全文
03-27 00:02 · AI,技术,HackerNews,AI
From zero to a RAG system: successes and failures
A few months ago I was tasked with creating an internal tool for the company's engineers: a Chat that used a local LLM. Nothing extraordinary so far. Then the requirements came in: it had to have a fast response, I insist... fast!, and... it also had to provide answers about every project the company has done throughout its entire history (almost a decade). They didn't want a traditional search engine, but a tool where you could ask questions in natural language and get answers with references to the original documents. With emphasis on providing information from OrcaFlex files (a simulation software for floating body dynamics, cables, etc., widely used in the offshore industry). It already seemed complex, but it was confirmed when I was given access to 1 TB of projects, mixed with technical documentation, reports, analyses, regulations, CSVs, etc. The emotional roller coaster had begun.
I'll tell you upfront that it was neither a quick nor easy process, and that's why I'd like to share it. From the first attempts, mistakes, to the final architecture that ended up in production. I also want to highlight that I had never done anything similar before and didn't know how a RAG worked either.
We'll go problem by problem, and the solution I applied to each one.
Problem 1: selecting the right technology
The first step was to define the stack.
I needed a local language model, without relying on external APIs, for confidentiality reasons. Ollama emerged as the most mature and easy-to-use option for running LLaMA models locally. I tried several embeddings, and nomic-embed-text
offered good performance and quality for technical documents.
Next was a RAG engine to orchestrate the document indexing process, embedding generation, vector database storage, and queries. Without it, no matter how fast the language model is, we couldn't retrieve relevant information from the documents. Think of it like a book's index: without it, you'd have to read the entire book to find the information you need. And with a good index, you can go straight to the right page. I'll call this process indexing for simplicity, although it's really a vectorization and indexing process.
After some research, I found a mature open source framework called LlamaIndex.
The language I'd use would be Python, I could list many reasons, but the most important one is that I feel comfortable and productive with it. Additionally, both Ollama and LlamaIndex have excellent Python SDKs.
I was ready to start building the software. I wrote my first scripts to run vector tests on the RAG system and do some query experiments. It worked really well with very little code. I thought it would be a project of a few weeks. I couldn't have been more wrong.
The next step was working with the actual documents. Hold on tight, it's going to be a bumpy ride!
Problem 2: the document chaos
My file source was a folder on Azure with a massive amount of technical documents: hundreds of gigabytes, thousands of files, various formats, with no organization or structure beyond the folder hierarchy. Every data engineer's dream (note the irony).
I cracked my knuckles, set the RAG output to save to disk, and launched my first script. LlamaIndex ended up overflowing my laptop's RAM within minutes, choking my OS until everything froze. I tried many configurations, caching systems, and other strategies, but at some point my machine always died.
After debugging, I discovered it was processing huge files that contributed nothing: videos, simulations, backup files... Documents that added nothing to a RAG system, but that LlamaIndex tried to process as if they were text. If a file weighed several gigabytes, the system tried to load it entirely into memory for processing, which was suicide.
I added a filtering system to the pipeline that excluded files by extension and by name patterns (simulation files, numerical results, etc.).
I also removed files that were expensive to process and didn't add value either, like CSVs, JSONs, among others. On the other hand, I converted PDF, DOCX, XLSX, PPTX, etc. files to plain text so LlamaIndex could process them without issues.
The result was a 54% reduction in the number of files to index. And of course, my RAM stopped exploding.
I could finally start indexing without fear.
Problem 3: indexing 451GB of documents without dying in the attempt
A RAG involves creating a vector index file containing document embeddings. Vectors are numerical representations of documents that allow measuring their similarity. LlamaIndex has a simple system you can configure with a couple of lines. You just point it to the directory and it takes care of storing all the information inside in JSON format. It's really convenient, works well, unless you're dealing with hundreds of gigabytes of documents. The system became unmanageable: every time the service restarted, it had to reprocess all documents from scratch, which could take days. Also, the default format is not optimal for larg
▸ 展开全文