Assemble the
Community
Part I built the infrastructure. Part II converted the people. Yet a recurring pattern ran through the study: organizations that did both – with individuals accelerating 10×+ on their own work – still couldn’t push organization-level productivity past roughly 50%.
Without the org changes that come next, it’s like upgrading a sedan to a sports car and then stopping at every stoplight: the speed is real, the organization just won’t let you use it.
Capturing the rest takes a new shape: requirements become the bottleneck, so the PM-to-engineer ratio breaks to an extreme; product-minded engineers organize into smaller autonomous squads; four experienced chiefs hold company-wide architecture, product, design, and security direction across them; and the org tree gets shorter and wider. The same building capacity then spreads beyond engineering into the wider organization. If your structure hasn’t changed, you haven’t really absorbed the first two parts.
“Adapt or die.” – Moneyball (2011)
The 50% Ceiling
Parts I and II make individual engineers dramatically faster. They do not, by themselves, make the organization move at the same rate. That gap between local velocity and organizational throughput is the central finding of Part III.
The pattern first surfaced at one of the study’s most advanced companies: agentic for more than a year, with 85–90% of production code generated by agents and a strong infrastructure layer. If any company should have translated local velocity directly into company output, it was this one. It could not.
The Ferrari at every stoplight
The analogy lands instantly with every CTO who has hit this wall. Replace an old, slow sedan with a sports car: the acceleration is sharper, the top speed is higher, and every measure of the vehicle improves. Then drive it along the same city route, with a red light every 200 meters. You arrive only slightly sooner. The same bottleneck appears in outside data. Across more than 400 companies, DX found median pull-request throughput rising 7.76% as AI adoption increased 65% – meaning much of the local acceleration was absorbed before it reached wider organizational output.
This is what happens when an organization installs AI-native engineering inside a pre-AI org shape. Some controls still protect quality and safety; others are habits inherited from the old speed. The answer is not to run every red light. It is to redesign the route: remove unnecessary stops, automate controls that can operate at machine speed, and create high-throughput lanes for agentic work. Product reviews, PM handoffs, review queues, month-long planning, human-paced status meetings, and legacy ownership boundaries all have to be reconsidered.
The lights were gone. He still stopped.
I spoke with an engineer who had been given a new project and an explicit mandate to work in AI mode. The organization had cleared away the normal approval gates and told him to move fast. Yet at each meaningful decision – architecture, scope, product tradeoffs – he still went looking for permission. The stoplights had been switched off; he kept stopping at the junctions.
Decades of product and software management taught the same sequence: define, align, review, approve – and only then build. That discipline made sense when code was expensive and rework was slow. It also trained engineers to seek permission before deciding, protect work once written, and treat discarded code as waste.
Agentic work asks for different instincts: make more reversible decisions, test them in software, and throw work away without ceremony when the evidence changes. Removing the gates is only half the transformation; people must learn to move without waiting for them.
When two AI-pilled engineers outpace your fifty.
The danger isn’t only your own ceiling. Remember the 46× concentration from Part I. Picture an R&D org of fifty that came late to AI, up against just two converted engineers at a competitor. On throughput alone, the two can outpace the fifty. But the arithmetic understates their advantage. At two people, there is barely an organization available to slow them down: no management layer, no PM handoff, and no queue of teams waiting on one another. If they also have the rest of the stack – deep domain expertise and the right architectural decisions underneath them – they are operating in full entrepreneur mode, while the fifty keep stopping at every junction. Your existing codebase is still a moat, but a bounded one: it makes an excellent spec, and two engineers like that can rebuild an equivalent baseline from it and then race past it. It’s a scary thought – which is exactly why you want those two working for you, not for someone else.
The organizations that broke through 50% rebuilt the workflow around the new speed. Part III is the map of that rebuild.
Product–engineering interface.
Reset PM ratios, planning, and handoffs.
Teams and direction.
Build small autonomous squads, held together by four company-wide chiefs.
Organizational structure.
Shorten reporting paths, widen spans, and extend building beyond engineering.
The PM-to-Engineer Ratio Inverts.
Historically, the typical software organization averaged roughly one PM for every six engineers. The ratio held because engineering was the slow side; product could define what to build faster than engineering could deliver it. That balance has now flipped. When a feature ships in a day instead of a sprint, the PM becomes the bottleneck. The question every CTO in the study has had to face is not whether to restructure the ratio, but in which direction.
Two valid directions. Customer proximity decides which.
Direction A – The ratio widens toward 1:10 or beyond. Engineers absorb product work. When engineers can credibly proxy the customer – because they are the customer in devtools or infrastructure, or because they know the domain deeply – the PM function compresses. Product work does not disappear; it moves into engineering. Each engineer becomes a mini-entrepreneur for a product surface: close to the customer, choosing what to build, shipping it, and learning from use. A PM, where one remains, covers a broader portfolio instead of feeding requirements to a single squad.
Direction B – Ratio inverts toward 1:1. PMs become hybrid PM/engineers. When engineers cannot proxy the customer (wrong geography, demographic, or domain), the PM cannot be removed; it has to scale up. PMs work in coding tools, ship clickable prototypes, and engineering’s job shifts from “translate the spec” to “harden the prototype the PM already built.” One late-stage converter is moving from 1:6 toward 1:2, every PM in Claude Code, on this logic:
The beta is the spec; the PR is the handoff.
Amazon made APIs the interface between teams. AI is now making working software the interface between functions. Designers can submit functioning interfaces; PMs can submit working features. People who did not know Git a year ago are now opening pull requests.
The handoff remains, but the translation loop disappears. Everyone can react to the beta instead of interpreting a specification. Engineering still owns architecture, quality, security, and production approval – but it begins with a working change.
The same handoff runs in reverse: production evidence, support patterns, and FDE learning can arrive as a tested PR rather than a ticket passed through the organization.
Cat Wu, who heads product for Claude Code, describes the convergence: “Our roles are blending together: designers ship code, engineers make product decisions, product managers build prototypes and evals.” The study’s working hypothesis on hiring tracks this:
Within twelve months, the company that has made one of these moves will not look like the company that made neither.
The same logic collapses the planning cadence.
The ratio is not the only casualty of fast shipping. The four-quarter OKR cycle, the six-month roadmap, the two-week sprint – all are artifacts of an era when a feature took a month. When it takes a day, planning six months out means planning around constraints that won’t exist by the time you ship; the loop has to shrink to match, with more experiments run and discarded and fewer long-dated commitments. And the model itself keeps moving underneath the plan – Anthropic’s internal benchmark puts it at “a roughly 41× jump in 16 months” in the length of task a frontier model can complete unaided, measured from Sonnet 3.5 to Opus 4.6. A planning cycle longer than the gap between capability jumps is planning around a world that no longer exists.
The Squad Is the Unit That Ships.
Product-minded ICs organize into small squads and command their own agents. Companies in the study converged on the unit, but not its size.
The range is narrower than the debate around it suggests: one to five people. Two to three is where most land. One AI-native company runs ~5 per vertical (2 backend + 2 AI engineers + an architect-shaped person). Two heads still beat one on most non-trivial problems, even when both are operating agents, and the social dynamics of pair-coding-with-agents stay healthier than full isolation.
The one-person squad is the extreme end of the same line, not a different structure. It remains contested: one AI-native startup in the study dismisses it as “romance – it’s not good enough yet.” But an advanced late-stage company runs five-to-six such squads in production. A single engineer owns products end-to-end, with a team of agents working nearly twenty-four hours a day. Each person operates as the mini-entrepreneur described earlier: they understand what the product needs to do, handle their own PM work, and command the agents. That configuration fits a domain where engineers can credibly proxy the customer; this company reports 7× velocity versus its pre-restructure baseline.
Junior engineers are back – conditionally.
Counter to the “junior is dead” narrative, multiple study companies report the opposite – on two conditions. The hiring filter is now AI proficiency and product sense, not years of experience; and the architectural spine is the precondition.
At scale, the picture inverts. One mid-stage SaaS company in the study, instrumenting its own engineering org, found exactly the opposite pattern: its L2s, the most recent grads, are faring worst with AI, while staff and principal engineers are the most engaged. The CPTO described it as “experience and taste – being able to make decisions faster with more context.” In a small AI-native organization, a strong architect nearby can carry more of the judgment load. In a larger company without that proximity, individual judgment matters more, and experience is what produces it. The filter still is not tenure, but at scale the taste that comes with experience compounds.
Four Chiefs at the Helm.
Small, autonomous squads ship with fewer handoffs. That freedom creates the speed. It also creates a new risk: every project can make sensible local decisions and still pull the company in a different direction. That is why the flatter organization needs a small set of company-wide functions to keep the system under control and help teams navigate at speed.
In the old structure, architects, senior managers, product reviews, design critiques, and security gates carried decisions across teams. Once those layers thin out, alignment no longer happens by default. A squad knows its product and an agent knows the repository and task in front of it. Neither naturally sees why another team chose a particular database, how two products should share identity, what customers have already been promised, or which details make the company’s brand recognizable. Four human roles hold that company-scale context:
Keeps architecture consistent across projects and current with the best technology elsewhere.
Aligns autonomous teams to one strategy and the customer demands that matter most.
Gives the company one visual and product voice, distinct and consistent across every surface.
Maintains one security boundary across models, agents, squads, and Part I’s risk surface.
The chiefs should be the company’s most experienced people in their respective fields and its strongest AI operators. Squads use agents inside one project; chiefs use them across many projects at once, carrying architecture, product, design, and security judgment from one team to another. Agents do the scanning, retrieval, comparison, and routine communication, so each chief can remain an individual or a very small team rather than growing into a large CTO office. Concentrating the role keeps standards consistent and reserves the company’s best human judgment for the decisions that matter most.
When every squad chooses its own route.
Fewer layers and routine gates mean more choices are made locally: databases and data-access layers, authentication, API patterns, UI components, state management, testing, observability, and external libraries. Model diversity widens the range of answers. Advanced teams should use several models: their strengths differ, independent opinions improve important decisions, and no company should depend entirely on one provider.
But models bring defaults of their own. A 2026 Findings of ACL paper, A Study of LLMs’ Preferences for Libraries and Programming Languages compared eight models and found a consistent popularity bias: they reached for familiar languages and libraries even when they were a poor fit. The operational point is broader: models can favor familiar technology over the best fit for the company. Several models are still the right practice, but their differing defaults add variation to the freedom squads already have, while shared training data can pull all of them toward familiar solutions. Across independent squads, locally reasonable choices can accumulate into incompatible stacks and duplicated infrastructure.
When all products look the same.
Visual design makes that convergence easiest to see. Models trained on much of the same product corpus reach for the same safe defaults: familiar component libraries, rounded cards, predictable layouts, and an increasingly familiar visual grammar. If the company has not set a visual voice of its own, the model fills the gap with the statistical average. The result can be polished and usable while still making one product difficult to distinguish from the next.
That voice must be set before it can be scaled: typography, color, composition, imagery, motion, and the conventions the company chooses to reject. The Chief Designer owns that graphical language and the human point of view behind it. Agents can then apply it consistently across hundreds of screens and generate variations without averaging the brand away. Three study companies independently made the same observation: “AI design is derivative – trained on Tailwind and Bootstrap, everything looks the same.” Without that human hand, the company can ship a technically distinct product that looks like every other AI-generated one.
One company memory.
A squad’s working context is intentionally local: requirements, repository, tests, and the customer problem in front of it. A chief’s context spans the shared architecture and stack, product strategy, security policy, design system, and customer commitments. It also records which code, data, prompts, and design assets can be reused, along with the exceptions accumulated over time. Maintaining that context is a real operating job: keep decisions consistent across repositories, preserve the company’s IP, and make what every team builds maintainable by the next one.
The chiefs use the strongest available models to search that history and compare outside approaches. On important questions they use more than one, so disagreements expose the defaults and assumptions described above. The models broaden the view; the final call rests on accumulated human experience.
Traffic control without stoplights.
Removing the old traffic lights is what gives autonomous squads their speed. But removing control altogether would produce chaos: standards diverge, local decisions collide, and risks surface only after the fact. Squad leaders are still expected to consult the appropriate chief before major architecture, product, design, and security decisions. But the chiefs do not depend on teams recognizing and escalating every issue. Each chief operates agents that scan plans, code, dependencies, design changes, security exceptions, and customer commitments across all active projects. The agents compare that stream with company standards and earlier decisions, giving the chiefs visibility no human could maintain alone. They can consume it asynchronously and step in when a local choice is becoming a company-wide one, while routine decisions remain with the squads and the work keeps moving.
The chief without the queue.
The same context gives the organization a way to reach the chiefs. In one AI-native company, the CEO built a personal agent connected to Slack, production, meeting transcripts, and monitoring; people speak with it as his representative before coming to him. Another advanced organization gives every product a bot that knows its conversations, code, and customer calls; its Head of Product uses one as the first pass on launch approval. These agents answer routine questions and carry accumulated context back into the squads, while decisions requiring judgment still reach the person.
The Engineering Org Tree Gets Shorter and Wider.
Once product-minded ICs can orchestrate agents, small squads can ship independently, and the four chiefs can hold direction across them, the management structure around the work can compress. As agents absorb status collection, task coordination, and routine handoffs, managers can cover more engineers and squads. Engineering management does not disappear, but its density and purpose change.
In a 25-person startup, that can remove the team-level manager entirely. In a 500-person engineering organization, it more often removes a layer, widens spans, and shifts the remaining managers from coordinating work toward people, portfolio, and cross-team judgment. The unit of change is management density: fewer managers per engineer, fewer hops from builder to executive, and more scope per manager.
The traditional engineering-manager role bundled people leadership with flow coordination: status meetings, sprint planning, capacity allocation, blocker removal, and cross-team negotiation. Agents reduce much of that coordination load, but not the need for coaching, performance management, staffing, governance, or decisions that span products and platforms. What compresses first is the layer whose primary job is translating status and coordinating handoffs between other layers.
More capacity can mean fewer people – or more software.
AI expands effective software capacity; it does not determine the size of the organization. Some people make the transition and multiply their scope; others do not. Some companies hold demand constant and use the gain to reduce headcount. Others widen the roadmap, enter new markets, and keep or expand the workforce. Demand for software is close to unlimited. The limiting factor is whether the company can identify valuable work and reorganize around the people who can deliver it.
Large organizations still need managers to develop people, make staffing and performance decisions, hold regulatory and customer commitments, and coordinate shared platforms across hundreds of engineers. But they need fewer management hops and fewer managers whose scope ends at a single small team. The role becomes broader and more judgment-heavy as its coordination machinery becomes increasingly automated.
Part of what makes the wider span workable is that the residual management job is itself being automated. The parts that scaled poorly with headcount – sitting in every meeting, writing status updates, prepping one-on-ones, chasing blockers, stitching together what happened across a dozen workstreams – are increasingly handled by AI: call and meeting transcriptions that summarize themselves, agents that draft status reports and surface risks straight from the team’s own activity, and assistants that turn a manager’s intent into the follow-ups. A manager who spends less time manufacturing visibility can hold far more people in view, which is how some organizations now run 15–25 reports where the old ceiling was six.
The bigger shift is in communication itself. AI is collapsing the cost of the visibility that used to cap how many people one person could track. Ramp runs its agents in shared, multiplayer sessions, so a team watches each other work in the open rather than waiting for a status sync; call and meeting transcripts harden into searchable corporate memory that outlives the conversation; and at one AI-native startup the CEO built a personalized agent of himself – his identity, wired into Slack, production, and monitoring – that people query before they go to him, in effect his standing representative. When visibility is ambient and a leader’s context is queryable on demand, the span a single person can hold stretches further still.
The management career path branches.
Engineering managers have built careers on the assumption that larger teams mean greater seniority. That equation is breaking. As spans widen and layers compress, the path separates into three credible directions. All three share a non-negotiable condition: the manager must become AI-pilled. Supervising people who use AI is not enough; leaders need to operate agents themselves, understand the new work firsthand, and recognize strong agentic performance.
(a) Broaden the management span. Manage several squads or a product area rather than one small team. AI supplies visibility and follow-through; the manager remains accountable for people, priorities, performance, and cross-team tradeoffs.
(b) Return to hands-on engineering and run agent squads. The former manager-of-four becomes a senior individual contributor orchestrating several agents, with comparable scope of impact but a different shape of work.
(c) Move into cross-squad leadership. Architecture, platform, product-surface, security, and other company-wide roles become more important as autonomous squads make more local decisions. These roles hold the context and judgment no single project can maintain.
The common thread is leverage, not team size. Large organizations will retain middle management, but fewer roles will exist primarily to relay status between layers. Managers who broaden their scope, deepen their technical contribution, or carry judgment across squads become more valuable; coordination-only roles become less durable.
When Building Leaves Engineering.
As the rest of the company awakens.
Engineering moved first, and its progression offers a template for what comes next. That does not mean turning people in revenue operations, finance, HR, or support into software engineers. It means giving each function agents and tailored internal applications that work across the systems where its work already lives. Instead of a person moving between the CRM, billing system, email, spreadsheets, and support queue, an agent assembles the context, executes the routine steps, and brings back the decisions that still need human judgment. As these systems improve, manual processes and thin SaaS point products begin to give way to agentic tools built around the company’s own data and operating model.
The study already shows this happening. One RevOps team is building sales-cycle analysis and deal-scoring software with coding agents and automation tools. Another company has a dedicated five-person GTM AI team building call briefings and other systems for sales and marketing. Elsewhere, companies have built their own candidate-screening tool for HR, an internal pricing application, help-desk and meeting-intelligence systems, and dashboards that combine sales, engineering, finance, and people data. These are not employees moonlighting as product engineers. They are functions rebuilding how their own work gets done.
The IDE-to-agent shift will repeat in business software. In engineering, assistance inside the existing interface was only the first step; the larger change came when an agent could take an outcome and operate the toolchain. A better sidebar inside the CRM or an AI extension inside Excel is still the IDE stage: useful, but waiting for a person to drive the application. The agentic shift begins when an agent reads the customer history, updates the record, triggers the next action, and works across CRM, email, billing, and support without waiting for a person to click through each application.
The chiefs’ context-management practice extends to strategy and marketing. The same operating model used to keep architecture, product, design, and security coherent can connect business operations with the market outside. Agents become active sensors across customer conversations, sales activity, support, finance, product usage, competitors, and public sources. They preserve not only what happened but why decisions were made, giving leadership and marketing one living body of context for roadmaps, positioning, and company strategy instead of another stack of fragmented reports.
The org tree gets shorter and wider across functions. The same tools that compress engineering management can strip routine coordination from every management role. Meetings become transcribed and searchable; agents track progress, surface blockers, prepare status, and follow up without another reporting cycle. Removing that work removes many of the human-paced stoplights between teams. Managers can carry wider spans and broader portfolios, reserving their time for coaching, judgment, and the decisions that actually need them.
AI Ops becomes a company-wide function. It no longer serves only engineering. It provides every department with tools, training, reusable patterns, and a common security and data foundation. It also gives the chiefs and functional leaders visibility into what is being built, which models and data it uses, and where local choices are beginning to diverge. That is how the company gains speed without ending up with shadow software, inconsistent controls, and ten different databases solving the same problem.
That moves the build-versus-buy boundary. As custom development gets cheaper, companies can build the agentic layer and the distinctive workflows around their data instead of buying a separate product for every process. The durable systems of record, regulated cores, and products with real networks still get bought; rigid point solutions and thin interfaces are easier to replace. The likely split is to buy the trusted core and build the workflow that makes the company different – which is why APIs, MCP support, and agent access are becoming procurement criteria.
Cat Wu observed the same pattern across Anthropic: “Our data science, finance, marketing, legal, and design teams picked up these tools on their own. The whole organization moves at the same speed instead of waiting on handoffs.”
Where You Stand, Where We Go from Here.
The thesis is easy to state and hard to live: agentic engineering is an infrastructure problem, then a people problem, then an organizational one – and most companies stall at that last step, the one with the real leverage. So rather than restate it, the more useful question is: where do you stand?
Read the clock. Then read the slope. The window has been short. The AI-enabled IDE made assistance ambient in 2023–24, but a human still drove every line; the real inflection came over Q4 2025 – Claude Code 2.0 in September, Opus 4.5 in November – when agents crossed from demo to dependable. How thin the ground was before that is measurable: METR’s randomized trial on early-2025 tools found experienced developers 19% slower with AI than without, while estimating they had been 20% faster. Re-run on the same developers late in 2025, the direction reversed – though they caution the newer figure is a weak signal, partly because too many of them now refused to work without AI at all. Almost everything in this playbook unfolded in the months around that moment, and where an organization sits today tracks closely to when it started, on two fronts at once: how far its tooling has gone, and how far its people have.
The AI-natives that were agentic from day one, plus the established companies that took a sharp turn in Q4 2025 – straight into autonomous agentic development, not merely AI-assisted IDEs. They are already past tooling and conversion, into the hard part: restructuring the org and tuning it on real measurement. Most of their staff sits in the Believers camp by now – moved there through training, recruiting, and replacement, and the obstacles their Legacy Guardians stood on have been cleared.
Where most organizations probably sit. They usually have some agentic projects underway – new products, AI in the product – but most of the org is still IDE-based, working on the infrastructure that would let agentic development scale. Not yet ready for a sharp turn: their teams still carry a large bloc of Legacy Guardians and Resistant who need convincing or replacement, and many feel constrained by their existing architecture and wary of losing control.
Some IDE tooling added, but every change still routes through mandated human review – by choice or by compliance. Earliest on both fronts, with the most ground to cover and the least time to cover it.
Timing created the lead. Operating choices set the slope.
The early advantage is real and compounding. The trajectory available to later movers is still a choice.
Timing created the early lead. High-agency companies took the turn first, and the systems and learning they accumulated are still compounding. But timing is not destiny. Later movers now inherit better tools, gathered experience, and a proven playbook, so they can traverse the early curve much faster than the pioneers did.
Operating choices determine the slope from here. Cautious adoption can produce meaningful local gains while leaving the operating model intact; on that path, the gap continues to compound. An operating-model reset – redrawing infrastructure, workflows, roles, and decision rights around agent-enabled work – can put a later mover onto the same accelerating curve and begin to narrow the gap. For later movers, the choice point is now.
The turns keep coming – and you need to be prepared.
The Q4-2025 turn was not the end of the story – it was the first in a series, and the ones after it land faster. Each one resets the structure: an agent that works unattended for an hour and one that works unattended for a day are not the same tool used longer – they call for different teams, different review, a different shape of organization around them. Every breakthrough sends a shockwave through the people and the org chart.
This is what makes the moment genuinely hard. We plan linearly – next quarter looks like this one, plus a little – and the capability does not. A structure built for a 5× world is behind in the 20× world that shows up while you’re still staffing for the 5×.
What might a year from now look like? Imagine you no longer open an editor on Monday morning – you brief a fleet of agents instead. Not the five or six that feel aggressive today, but hundreds, spread across features, fixes, tests, migrations, research, and review – most of them working through the night and the weekend without a prompt from anyone. Your job is to decide what gets built and to confirm that what comes back is right. Meanwhile, the codebase you counted on as a moat protects you a little less each quarter, now that an equivalent can be rebuilt from scratch in a fraction of the time. None of this needs a breakthrough you can’t already see coming – it’s just today’s curve, drawn out twelve months.
It won’t only be more agents – it will be faster ones. Today’s agentic loop will feel to you like dial-up speed: you hand off a task and wait minutes, sometimes longer, for it to come back – the modern equivalent of a 300-baud modem, watching a page of text crawl onto the screen a line at a time. That latency is collapsing the way bandwidth did.
The cycles that cost you minutes today – a build, a test run, an agent reworking a change – compress toward a snap and return not just quicker but more accurately, with fewer wrong turns to unwind. When the short loops that need your input close in seconds rather than minutes, interactive work stops being a chain of hand-offs and starts to feel like a live conversation. Longer loops move in the other direction: agents remain on a problem for hours or days without pulling you back in. The one thing that doesn’t speed up is your judgment about what to build and whether the answer is right – which is precisely why that becomes the whole job.
And it spreads past engineering. Revenue operations, finance, HR, support, and legal are already building their own agents and internal workflows. R&D moved first – it had the longest history with the tools and the strongest incentive to automate itself – but the rest of the company does not have to discover the path from scratch. The playbook for tooling, conversion, cross-functional context, and organizational change can now be carried over.
The only real certainty is that it doesn’t stop. A finished reorganization will not be enough; the discipline is to stay close to the edge and see each shift early. The agentic awakening is already moving through the industry. Those who see it early will define the next era of software. Everyone else will wake up inside it.