On September 2, a discussion began in LLVM about adding "AGENTS.md," a guidance file for AI coding agents, to the top level of the repository that manages its source code. The draft AGENTS.md proposed by Nick Desaulniers is just two lines long and prompts agents to read the project's existing AI usage policy and coding standards. The goal is to get the rules to the people bringing in generated output before reviewers have to repeat the same warnings. Developers, however, have raised a stream of doubts about whether a shared document would be effective and how much upkeep it would require. LLVM already permits AI assistance, so why is this small addition provoking debate?
What the two-line draft would change
The two lines in draft PR #220659 boil down to one instruction: check the AI usage policy and coding standards before contributing. As of the check on September 4, Japan time, it was still a Draft, and no decision to adopt it had been made. Rather than copying the text of the rules, it refers to documents that already exist in the repository.
AGENTS.md is a Markdown file that tells AI agents things like build and test procedures and how to write code. It supplements the human-oriented README and gives AI a consistent place to look before working. However, simply putting text there has no technical power to prohibit posting or file operations.
In the RFC, Desaulniers noted that LLVM is moving its documentation from reStructuredText to Markdown, and explained that this would let agents refer directly to the existing documents humans read. The idea is to avoid maintaining separate rules for people and for AI. On the other hand, AGENTS.md and CLAUDE.md are currently excluded from Git tracking, so sharing the file would also require changing the ignore settings, he said.
The scope of the proposal is narrow. Per-subproject AGENTS.md files, SKILLS.md, and model-specific configuration directories are out of scope. It also mentions referencing AGENTS.md from CLAUDE.md, but that too is only at the proposal stage. This is not a change to the current conditions for AI use; it is a change in whether the project provides an entry point that guides agents to those conditions.
The review time LLVM is trying to protect
LLVM's current AI tool policy requires a person to read and verify AI-generated code or text before asking others to review it. Contributors must be able to answer questions about the content, and if they used a substantial amount of generated material, they must disclose it, for example in the PR description. Posting automated review results without human approval is also not allowed. The person ultimately responsible for generated output is the contributor.
This line is meant to prevent contributions that push the burden of verification onto maintainers. The policy asks contributors to provide value commensurate with the review time spent, and seeks to avoid a situation where maintainers receive unchecked output and have to hunt for problems. It also places weight on newcomers learning and becoming future maintainers. For that reason, using AI to resolve good first issues, which are set aside for learning, is prohibited. There is an exception for the approved Bazel-fixer bot. It is not a blanket ban.
For example, if an AI produces a fix but the contributor cannot explain why the change was made, the reviewer has to start by reconstructing the intent. Telling people the rules in advance is meant to head off this rework at the entrance. But an agent having read the rules is not the same as a human having understood the output. Even if AGENTS.md is adopted, someone still has to take responsibility for checking before posting.
Can shared guidance be separated from personal instructions?
In the discussion, Jonas Devlieghere said his own LLVM development skills have been useful, and supported sharing to reduce the burden of everyone maintaining the same guidance separately. When he used the general LLVM standards in LLDB, he said, the agent did not inappropriately rewrite code to conform to them. He admitted, though, that he could not tell whether extra computation or searching was happening internally.
People who already use their own instructions face a different concern. Reid Kleckner explained that he links an instruction file from a private repository into each working copy to convey personal settings such as the build directory. If an official AGENTS.md were placed in the same location, he would need to decide how to add to or override it with his own settings. Standardization could force people to rearrange existing working environments.
Fangrui Song (MaskRay) objected, citing the nature of settings kept in a shared repository. His argument is that settings whose results can be verified, like .clang-format, differ from instructions that shape how an individual works. Formatting and static checks are already enforced in CI, and the AI policy defines obligations for humans. He holds that it must be separately shown what adding instructions for agents would improve.
Differences among subprojects cannot be ignored either. LLVM's coding standards give priority to the style of existing code, and state that libc++ departs from the general standards to match the C++ standard. Even within one repository, the same rules cannot be applied everywhere.
An experience shared by Cullen Rhodes makes the problem of instruction lifespan concrete. When working on the compiler's internal GlobalISel and generic MIR, he said, the AI wrote code that redundantly checked the number and kinds of arguments already guaranteed by the definition or the immediate caller. He added an instruction to avoid such unnecessary checks, but hesitates to propose it for a shared file. It depends on his own agent, its effect is hard to verify, and it may become unnecessary as models improve. This illustrates the difficulty of promoting an instruction found through individual trial and error into a rule everyone must maintain.
What research measured, and what it didn't: responsibility
Aaron Ballman raised doubts about effectiveness and cited a paper by Thibaud Gloaguen et al., published on arXiv. Version 2, released February 12 and revised June 23, examined four models and their supported agents using 138 tasks from 12 repositories in CTXbench and 300 tasks from 11 Python repositories in SWE-bench Lite. It is not an experiment that evaluated LLVM.
In the latest version's results, the differences in success rate between no instructions and LLM-generated instructions, and between no instructions and developer-provided instructions, were both not statistically significant. Developer-provided instructions significantly outperformed LLM-generated ones, but this does not confirm an improvement over having no instructions. Meanwhile, LLM-generated instructions raised average inference cost by 20% on SWE-bench and 23% on CTXbench. The instructions were generally followed, and testing and exploration increased.
These results call for caution toward the expectation that more instructions lead to better work. However, success rate here means the share of fixes that pass all tests. It does not measure whether contributors can explain their changes, whether unauthorized posting is avoided, or whether maintainers' review time falls. Kleckner also pointed out that behaviors such as not commenting on GitHub on its own using the user's credentials do not show up in success rates or token counts. If LLVM is to decide on adoption, it needs to check the correctness of fixes and posting behavior separately.
How far can the review burden be reduced?
The RFC also drew a proposal to follow rule links only when the relevant conditions apply. MattPD pointed in that direction, and Louis Dionne said he would support a minimal setup that directs agents to libc++'s tests and contribution documents. Neither is an agreed design, but the idea of splitting entry points by where the work happens, rather than making everyone read the same long instructions, is visible.
Based on the draft and current rules, the division of roles looks like this.
| Function | Current owner | AGENTS.md role / limits |
|---|---|---|
| Pointing to references | AI Tool Policy, Coding Standards, subproject documents | Can direct agents from the top level to existing rules. Does not automatically resolve per-location exceptions or personal settings. |
| Mechanical checks | CI, .clang-format, static analysis | Can point to checks that should be run, but does not replace the checks themselves. The current draft is limited to references to existing rules. |
| Contributor accountability | Contributors who read generated output, and the current AI policy requiring human approval | Serves as an entry point that conveys responsibility. Cannot enforce that the contributor understood, can answer questions, and approved. |
Pointing to references, mechanical checks, and contributor accountability each need to be verified separately.
From this distinction, it is also hard to judge practicality by the brevity of the guidance file alone. Even if the text is two lines, processing increases if agents are made to read everything it links to every time. Conversely, if brevity is prioritized and agents miss the rules for the area they are working in, maintainers end up repeating the same warnings later. The idea of referring to documents depending on conditions aims to balance the amount read against the accuracy of guidance.
The criteria for future decisions are whether agents reach the right subproject's documents, whether repeated warnings about rule violations decrease, and whether contributors understand and approve their changes. If LLVM can confirm these behaviors, it can weigh the value of adoption by considering both the effort on the AI-using side and the review time on the receiving side.
