In this video, Cursor engineer Lauren Tan (@poteto) explains how they got to the point of merging more than 1,000 PRs a month using AI agents. Tan's core argument is that the biggest problem in AI coding is not code generation itself but how to verify and trust what the agent produces. To do that, you need to build verification automation, feature maps, evals for agents, and strong CI constraints and architecture together. In the end, Tan says, a developer should become not someone who writes every line of code directly, but someone who designs the system so that agents can do good work.
1. Lauren Tan's Background and Today's Topic
Lauren Tan introduces themself as an engineer who goes by poteto on Twitter and joined Cursor about five months ago. Before that, Tan worked on the React Compiler on Meta's React team and still contributes somewhat to the core team and open source. Before Meta, Tan spent about two years at Netflix as a tech lead and engineering manager.
Having moved back and forth between individual contributor (IC) and manager roles, Tan finds that managing people and managing AI agents are surprisingly similar. That connection is the heart of this conversation.
"There are so many parallels between management skills and the way you manage agents."
The host mentions that they will also cover Grokbot, which Cursor recently launched. Grokbot is described as a product that gives multiple agents their own roles and identities and coordinates them. There was a brief hiccup with screen sharing and muting, and Lauren says with a laugh:
"It's 2026 and I still don't know how to use Zoom."
2. If You Can't Trust the Agent, the Human Becomes the Bottleneck
The biggest problem Lauren felt when coding with AI agents is simple.
"How can I trust this?"
An engineer who has written code by hand for a long time has their own standards for telling good design from bad, plus lessons learned from past failures. But agents often guess confidently, or call something the "smoking gun" when it isn't the real cause. When this keeps happening, trust breaks down, and you end up unable to make full use of agents.
Tan compares this to a manager micromanaging. If a manager doesn't trust their team, they have to keep watching over everyone's shoulder to make sure nobody ships bugs. Agents are the same. If you can't trust the output of even one agent, you can't run 100 agents at once.
"If you can't trust the output of a single agent, you can't spin up 100 of them."
At first, you have to personally watch every output and tool call of one or two agents and adjust your prompts. At this stage the human plays the role of verifier, so parallelization is nearly impossible. When the agent implements something, a person has to open the app, reproduce the error, and copy screenshots and console logs back to the agent.
"If you don't have verification, you are the verifier. Which means you are the bottleneck."
By contrast, having built up enough trust, Tan now runs agents that auto-merge PRs. One morning about 20 PRs had already been merged, and Tan reviewed them after the fact on the main branch. The results, Tan says, were actually fine.
"I woke up today and about 20 PRs had already been merged. I reviewed them on main, and they were good."
3. The Foundation of the Productivity Explosion Is Verification, Not Code Generation
Tan says their productivity was low during the first month at Cursor while getting to know the codebase. But as Tan gradually learned how to trust agents, that productivity shot up. Last month Tan shipped 1,000 PRs, and this month, with only 12 days gone, has already merged nearly 800.
These numbers aren't meant as bragging; they illustrate how the amount of work Tan can handle grew as the level of trust rose. Tan also acknowledges that many people may doubt whether "that much code is really good code."
The most important capability Lauren points to is verification. Verification here means more than just running tests. It refers to the agent's ability to actually run the application, operate it like a user, capture performance traces, analyze heap snapshots, and open the iOS simulator to confirm the problem is really fixed.
"Verification is the agent's ability to actually run the code, take CPU traces or heap snapshots, open the iOS simulator, and test it exactly the way a user would see it."
Verification doesn't perfectly guarantee code quality. But it at least lets you confirm the code works correctly. Lauren stresses that this is a "huge step forward" toward trusting agents.
"It's not a guarantee that it writes good code. But at least it lets it write code that works correctly."
4. Teaching Agents How to Use the Product with a Feature Map
Tan shares an early experience improving Cursor's agent window, the area known internally by the codename Glass. Tan was originally supposed to join the cloud agents team, but because of their React experience was asked to help with the agent window, which was about to launch.
The problem was that there was only about a week until launch, and Tan barely knew the new codebase. Tan had to capture performance traces in Chrome DevTools by hand and interpret flame graphs. Even when shown screenshots of traces, the agent would plausibly guess at a cause, and often it wasn't the real one.
"The agent would confidently say, 'It's probably this,' but when I fixed it, that wasn't actually the problem."
One of the early tools Tan built to solve this was a verification skill called Control Glass. The skill helps the agent run the application, collect performance traces, and control the program. Tan explains that for Electron, web, or iOS apps you can build a similar control system through things like the Chrome DevTools Protocol or Apple's simulator tooling.
But being able to run the application doesn't mean the agent understands the product. Even when users reported "the left sidebar is slow" or "the PR tab doesn't work," the agent didn't even know how to navigate to that feature. It would spin up a dev build and then wander around clicking here and there.
So Tan created a feature map. A feature map is a document that tells the agent how to reach each feature of the product. It contains the entry paths from the user's point of view, keyboard shortcuts, UI layout, and the DOM selectors and attributes needed for automated control.
"A feature map teaches the agent how to get to every feature we have."
Thanks to the feature map, even very thin user reports become workable. Cursor's internal feedback channel apparently gets plenty of vague reports that are just a single screenshot and a question mark. Without a feature map the agent can do almost nothing with these, but with one it can get context on which screen and feature to look at.
"We get a lot of reports that are just a screenshot and a question mark. Without a feature map, the agent can only say, 'I have no idea.'"
The P-stack plugin Tan built includes a create verification skill that helps bootstrap this kind of system automatically, and a maintain verification skill that keeps it updated as the product changes.
5. P-stack Is Agent-Operations Knowledge Built by Watching Failures
The P in P-stack is a playful name combining Tan's nickname poteto with "sack." It was inspired by Y Combinator CEO Garry Tan's G-stack, and Lauren Tan explains with a laugh that although they share a last name, the two aren't related.
"The P in P-stack is for poteto sack."
Building a plugin was never the original plan. While closely observing the ways agents fail, Tan began turning each recurring failure into a skill. For example, after seeing the agent confidently guess at a cause without properly reading the relevant code, Tan wrote guidelines like "don't guess—search and check the code" and "make active use of subagents."
"Every time I saw a failure mode, I'd say, 'OK, let's make this a skill.' Don't hallucinate, go find the actual code, stop guessing."
Tan likens this to onboarding a great newly hired engineer. Someone who is an excellent coder but has no business context needs to be taught the context and procedures the job requires. Agent skills are essentially Markdown documents, but inside them you can put rules, product context, troubleshooting procedures, and expected behavior.
Because an LLM is a model that predicts the next token, giving it high-quality context from the start makes it more likely to reason in a better direction. Lauren notes that some people describe this as "steering the agent into a better latent space," and explains that while it sounds complicated, it ultimately means good input leads to a good way of working.
6. How to Evaluate and Improve Skills
The host asks how Tan maintains skills when the product keeps changing, and how to judge whether the verification system is trustworthy enough. Tan's answer is evals.
You can think of evals as unit tests for agent skills. You don't necessarily need a special framework; you can build them yourself at whatever level of rigor you want. P-stack's eval playbook tests skills in a fairly systematic way.
In Tan's approach, a main orchestrating agent first creates a rubric for evaluating the skill's expected behavior. It then has multiple subagents run and do the work in separate directories, with care taken even over the directory names so the agents don't realize they're being evaluated. That's because agents can behave differently once they know they're being evaluated.
"Agents can notice they're being evaluated, and when they notice, they change their behavior."
Cursor's support for multiple models is also an advantage. You can evaluate the same skill across several models to see how well it works on each. Every time Tan modifies a skill, Tan runs the evals to confirm the expected results really come out.
Evals can also be scored, and you can reduce bias by having one model judge the eval results produced by another. Using Cursor's looping feature, you can even run a task like "keep evaluating and improving until every item scores 10 out of 10." Tan explains that the Control Glass skill was iteratively improved this way too.
But some parts of skill maintenance can't be solved by automation alone. In the early stages, a person needs to pay close attention to the agent's tool calls, the code it reads, and its reasoning process, and catch where it fails.
"Maintaining good skills takes a lot of observation and taste. You have to become a kind of really good backseat driver."
Rather than passively watching the agent, Tan advises actively analyzing its behavior and finding flaws early on. Just as you'd think "why did they do it this way?" while pair programming with a colleague, you trace the agent's decisions, look for recurring patterns, and turn them into rules or skills.
7. Scaling from Local Verification to Cloud Automation
When you first build a verification system, Lauren recommends starting in the local environment. Locally, a person can directly observe how the agent launches the application, which APIs it calls, and how it manipulates the screen.
"If you're building a verification skill, start locally. That way you can observe what the agent is actually doing."
Tan, however, now uses cloud agents heavily. If you set up the environment and verification skills well, cloud agents can become not just a tool that helps one developer but an automation foundation that strengthens the whole team and company.
An example is Benny, an internal agent that handles bug reports. Benny takes incoming bug reports, spins up a separate machine in the cloud, launches Cursor, and tries to reproduce the bug using the same control skills. Going beyond just saying "I found the problem," it can also check whether the bug has already been fixed on the current main branch.
"Benny reproduced the bug and confirmed it had already been fixed on main. Then all I have to do is release a new build."
This automation saves a person an hour of checking "is this really already fixed?" But Lauren warns against jumping straight to running hundreds or thousands of cloud agents without trust. Mass execution without a verification system can just waste tokens and inflate costs.
The host summarizes Tan's journey like this:
- First, build skills locally that verify whether the agent actually produces correct code.
- Once trust is established, scale up in the cloud so more agents handle signals like bug reports on their own.
- Finally, you can reach the point of auto-merging PRs.
Lauren agrees with this summary and says there are no shortcuts in this process.
"This is about your personal level of trust in agents, so you can't just leap from here to there in one jump."
8. In the AI Era, Refactoring and Rewrites Need Rethinking
Lauren says the way we view refactoring and rewrites may also change in the era of AI agents. Traditionally, engineers advise "don't casually rewrite an existing system." Joining a new company and wanting to tear out all the old code you see is a common urge, but it's generally considered dangerous.
Tan argues that, depending on the situation, there are reasons to seriously consider a rewrite. In particular, Tan points out that problems once faced only by big tech companies are now becoming problems for many organizations. When tens of thousands of developers contribute to one giant monorepo, as at Meta, code quality is hard to keep uniformly high even with many excellent engineers.
"Before AI slop code, there was human slop code."
The infrastructure and development rules at large companies are generally designed so that even the least experienced or most error-prone member of the team can't cause a major incident. Frameworks, conventions, permission restrictions, and guardrails are safety mechanisms of the kind that stop an intern from accidentally deleting the production database.
An existing codebase with these mechanisms well in place—a brownfield application—can actually be a good environment for AI agents. Agents operate within existing constraints and conventions, so they're less likely to cause a big incident.
On the other hand, a newly started greenfield application is both the biggest opportunity and a dangerous area. If you prototype quickly, as with Grokbot, and "vibe code" without humans reading much of the code, the agent will take the fastest, easiest path every time. Over time, the codebase piles up short-term fixes and becomes hard to control.
"In a vibe-coded app with no guardrails at all, the agent just solves problems the easiest way. Then the codebase gradually spins out of control."
So you need strong structure and constraints from the very beginning. Tan says it took more than 600 PRs to restructure Grokbot onto a new architecture, and that this work built enough trust that reading all the code personally is no longer necessary.
"I've gotten to the point where I hardly look at the code anymore. I'm not saying that to sell tokens—I mean I did an enormous amount of work to get to a state where I don't have to look at the code."
9. The Dune Architecture and Enforced Guardrails
Grokbot's new architecture goes by the internal codename Dune. Lauren likens it to something like Next.js for Electron apps, and describes it as a framework designed specifically so that agents can write code well.
CI in this environment is quite strict. Wherever agents frequently fail or show bad habits, Tan blocks it with hard rules as much as possible.
A prime example is the ban on React's useEffect. Tan sees useEffect as something in React that easily leads to big mistakes. If you use it in Dune, CI fails.
Interestingly, Tan also banned code comments. That's because agents often recorded temporary review feedback or past conversations in comments as if they were permanent rules, without understanding the context.
"Agents leave comments like 'Lauren said never to do this.' But I was asking them to fix a specific part of that PR, not creating a global rule."
Tan's basic principle is clear.
"Anything agents are bad at, I ban."
In Electron apps in particular, separating the renderer process from the main process is important. The renderer thread draws the UI, so to maintain 60fps each frame must be processed within about 16 milliseconds. If heavy computation or lots of I/O gets mixed into the renderer, you get dropped frames, long tasks, and jank.
Grokbot splits directories like electron-main and electron-renderer, and CI checks the dependency graph to enforce that code from one side isn't wrongly imported into the other. In other words, performance problems aren't left solely to reviewers' attentiveness.
"If heavy computation or I/O runs in the renderer, you get long tasks over 16 milliseconds, and eventually the UI stutters."
Lauren describes the protective system for a good codebase in several layers.
- Architecture and directory structure: Gather each feature's code in one place, and make the most natural implementation path also the correct one.
- Static analysis and CI: Block bad imports, banned patterns, performance risks, and the like at build time.
- Lint and compiler diagnostics: Automatically detect recurring bad patterns.
- Rules, skills, and Bugbot: Give agents guidance, and use code review tools to find additional problems.
- Clear feature conventions: Make existing patterns easy for agents to imitate, reducing the cost of searching and guessing.
Tan notes in particular that agents have a strong tendency to copy existing code and take the fast path, so you need to design things so that "the shortest path is the best path."
"Agents love shortcuts. So make the shortcut itself the best solution."
If you keep feature-level code together in one directory, the agent can go into that folder and finish most of the work without repeatedly searching globally. This also means designing for "the least capable agent."
10. Don't Rely Only on Human Code Review—Encode the Rules
Lauren stresses that rules, skills, and Bugbot alone aren't enough. These are useful, but they're soft constraints that agents may forget or not follow consistently. You need mechanisms that actually fail on violation, such as strong CI, type systems, and architectural constraints.
"If all you have is rules and skills, Bugbot, and a style guide, it's only a matter of time before the codebase becomes a mess."
Tan thinks the reason Rust is getting attention again is similar. The Rust compiler and borrow checker provide strong constraints. Of course you still have to avoid writing unsafe code, but at least once the code compiles, you get a certain level of confidence.
"When the compiler enforces things strongly, people don't have to go check everything one by one."
Conversely, the worst state is "code review hell," where a human reviewer has to personally catch every invariant and rule on every PR. If you find yourself making the same comment over and over in reviews, Tan sees that as a code smell that should be solved systemically.
"If you have to comment 'you can't do it this way' on every PR, you should think about how to turn that into a lint rule or a CI failure."
In other words, problems people keep finding should be turned into one of the following:
- Detect them automatically with a lint rule.
- Enforce them as a CI failure condition.
- Change the architecture so that the mistake itself can't happen.
11. PR Size and Why Git History Matters
Tan's PRs range from a few dozen lines to several hundred, and sometimes several thousand. There's no strict size limit, but Tan encourages agents to split large work into multiple PRs.
The reason is that Git history is important context. When a PR clearly describes a small unit of change, it's easy to revert if a bug appears, and easy to trace which change caused the problem.
"Git history is an incredibly rich source of context."
If you put every change into a single PR tens of thousands of lines long, it's hard to know what changed and why. Atomic PRs matter not only for human understanding but also so that agents and future automation systems can make use of the change history.
12. Token Cost Is a Question of ROI, Not Spending
Viewers ask whether Tan's approach is only possible at an AI company where tokens are close to unlimited. Tan candidly admits that working at an AI lab means few constraints on token usage, and says they can't tell everyone to copy this approach exactly.
Still, Tan thinks token usage shouldn't be seen simply as a cost—especially for engineering leaders and startup founders, it should be judged in terms of ROI (return on investment). Refactoring the codebase up front and building the verification system and CI constraints takes a lot of tokens. But if you assume a world where agents will write most of the code going forward, that initial cost can bring large efficiencies in the long run.
"It's true you'll spend a lot of tokens in the early stages. But if we're heading toward a world where agents write all the code, it's a question of return on investment."
Tan says building such a system with people alone could take years. An engineer's own time and salary are costs too, so you have to compare "hiring more people" against "spending tokens to build a codebase where even the simplest agent can work well."
"If I'd done all this refactoring and verification myself in the pre-agent era, it would have taken forever."
The results don't stop at individual productivity. PMs, designers, and engineers unfamiliar with the codebase can also contribute code in a sustainable way. Tan believes tokens are never free, but that investing in them properly has a big effect in growing the capability of the whole team.
13. An Organization Where Non-Developers Build Products with Agents Too
Finally, the host asks how PMs and other roles keep up when engineers ship features this fast. Lauren answers that Grokbot plays a very big role here.
Cursor's existing IDE, CLI, and agent window are powerful, but they were fundamentally developer-centric tools. Non-developers could use them for knowledge work, but they weren't really comfortable interfaces suited to how those people work.
Grokbot, by contrast, Tan explains, gives people outside technical roles an experience like "the moment you first use Cursor." In a familiar, iMessage-like interface, you can give agents names and coordinate multiple agents like teammates.
"Grokbot is, I think, the Cursor moment for people outside of tech."
For example, a PM can have an agent that summarizes what Tan worked on overnight. Designers and PMs can also find a bug, fix it themselves, and ask, "I fixed this—could you review it?" Tan says that when actually reviewing such PRs, sometimes they're perfect.
"A PM will say, 'There was a bug so I fixed it. Can you take a look?' and sometimes I review it and say, 'Great, approved.'"
This isn't just because AI writes good code; it's because the Dune architecture and strict constraints keep even non-experts' changes within safe paths. As a result, designers and PMs can ship features themselves, and the Grokbot team can move much faster. 🚀
14. Closing
Lauren Tan's message is that the key to boosting AI agent productivity isn't longer prompts or blindly running more agents. The key is building a verifiable execution environment, feature maps that teach how to use the product, skills that are evaluated repeatedly, and architecture and CI with enough enforcement power to replace human review.
"What matters is creating an environment where agents can do well."
Ultimately, the developer's role is shifting from writing and reviewing all the code directly toward designing systems where agents—and even non-developers—can contribute safely. Lauren wraps up the session by inviting anyone with further questions to reach out via Twitter DM.
