H
Howardism
Plate IISynthesesHOWARDISM

The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence

PublishedJuly 16, 2026FiledEssayDomainSynthesesReading8 minSourceAI-synthesised

Resolves acceleration-whiplash's open question: the Faros-vs-CMU under-review 'divergence' is mostly measurement artifact — a vendor's adoption-depth *delta in unreviewed-PR count* over enterprise all-PRs vs a non-vendor *calendar-time share* of unreviewed agent PRs in open source — and both fit one story: total unreviewed output rises with volume while the share of agent PRs merged unchecked falls as teams learn risk-triage; the volume-concentration clause is supported (median per-project no-review ≈0%, pooled >50%; triage by PR type), and the residual disagreement is a forecast — whether triage discipline survives agentic authoring crossing from <1% to double digits

Illustration for The Under-Review Divergence: Faros's Widening Crisis vs. CMU's Convergence

Question#

Faros reads under-review as a widening crisis; CMU's non-vendor GitHub telemetry finds the agent no-review rate converging toward the human baseline (>50%→~14%) as orgs learn to review agent code. Is the divergence real (enterprise vs open-source populations, adoption-depth cross-section vs calendar-time trend) or does the whiplash's under-review pressure only surface where PR volume is highest? (from Acceleration Whiplash open questions)

Short answer#

The divergence is mostly not real — the two studies measure different quantities, on different populations, along different time axes, about different authorship units, and both are consistent with a single underlying story: the total volume of under-reviewed output rises as AI floods the pipeline (Faros), while the share of agent PRs merged unchecked falls as teams learn to triage review by risk (CMU). The question's second clause is supported: under-review is a concentrated, volume-and-risk-triaged behavior, not a uniform erosion. What genuinely remains contested is not the data but the forecast — whether the learned triage discipline survives agentic authoring crossing from today's <1% of PRs toward double digits. Both sources themselves flag exactly that as the untested break point.

Four axes of non-comparability#

The headline numbers are not the same measurement pointed at the same thing. Lining them up:

AxisFaros (Acceleration Whiplash, vendor-claim)CMU (Review as the Control Point, empirical)
Metric+31.3% increase in PRs merged with no review — a delta in count/incidence ("the most urgent finding")No-review rate among agent PRs: >50% (mid-2025) → ~12% (Feb 2026) vs a stable ~14% human baseline — a level, over time
PopulationEnterprise telemetry: 22,000 developers, 4,000 teams (Faros platform customers)Public GitHub: 2,860 ≥10-star repos that already contain agent PRs (2.5M+ PRs; no non-adopter control)
Time axisAdoption-depth cross-section: low- vs high-AI-adoption periods within-companyCalendar-time trend: full PR histories Jan 2020–Feb 2026, re-scraped
Authorship unitAll PRs in an environment where AI-assisted human-driven authoring dominates (agentic authoring <1% of PRs, per AI as Primary Author)Specifically agent-associated PRs vs human PRs

The arithmetic reconciliation is direct: a falling rate and a rising count coexist whenever volume grows fast enough. Faros's own throughput numbers supply the volume — +210% code-specific tasks per team, +16.2% PR merge rate, +33.7% task throughput (Acceleration Whiplash). If merged-PR volume roughly doubles while the unreviewed share declines from its early peak, the absolute number of unreviewed merges still climbs — Faros's +31.3% and CMU's >50%→~12% can both be literally true of their respective populations at once.

The time-axis difference does the rest. A cross-section over adoption depth confounds cohort learning: at any calendar moment, the deepest adopters are also processing the most AI output, so they show the most under-review — even if every cohort's rate is falling over calendar time as it learns, which is exactly what CMU's longitudinal series shows. Faros's comparison cannot see the learning curve; CMU's cannot see enterprise adoption depth. (Faros's own 2025→2026 comparison is explicitly "directional only — independent cross-sections, not a longitudinal panel" — Acceleration Whiplash.)

The volume-concentration clause: supported#

The question's alternative — "does the under-review pressure only surface where PR volume is highest?" — has direct support in CMU's appendix detail (3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse):

  • The median per-project no-review rate is near 0% for both author groups in every month, while the pooled agent rate started >50%. Merging without review is a minority behavior concentrated in a subset of projects — and the pooled series is dominated by whichever repos push the most agent PRs. The "crisis-level" early number was never a uniform property of agent-adopting projects; it was the signature of high-agent-volume outliers.
  • Projects triage by perceived risk, not indiscriminately. Agent no-review rates vary widely by PR type — test 69%, refactor 41%, bug fix 25% — while human rates are nearly flat (8–14%). Where review is skipped, it is skipped selectively on low-stakes changes.
  • The convergence itself is a volume-adaptation story: "an initial willingness to merge agent output unchecked gives way to reviewing it much as human PRs are reviewed" — i.e., the pressure surfaced early where agent volume was highest, then process adapted. This is CMU's third moderator (process adaptation) operating in the wild (Review as the Control Point).

The kicker is that Faros prescribes the very behavior CMU observes emerging organically. Faros's remediation #2 concedes that "every PR must have a human review… will break under the weight of volume" and prescribes risk-tiered gating — human eyes for high-stakes paths, agent review for lower-risk areas, no PR ungated (AI Engineering Report 2026: The Acceleration Whiplash). That is the triage-by-risk pattern in CMU's PR-type breakdown. On the mechanism, the vendor and the non-vendor source agree; only their headlines diverge — and the headline divergence tracks their incentives and framings (Telemetry vs. Survey Measurement: Faros's "widening crisis" sells the visibility platform; CMU's thesis needs review to be a steerable control point).

What genuinely remains open#

  1. The volume test hasn't happened. Faros's dataset has agentic authoring at <1% of PRs and warns that removing the human from the loop puts "an order of magnitude greater" pressure on every metric (AI as Primary Author); CMU's own open question asks whether convergence holds "as agentic authoring crosses from <1% of PRs toward double digits" (Review as the Control Point). Both sources agree the observed regimes don't test the regime that matters next. The convergence is evidence teams can adapt at current volume — not that triage discipline survives 10× more.
  2. Even the direction of oversight change is contested among studies. CMU's project-level convergence runs opposite to Yu et al.'s within-reviewer habituation (oversight weakening as approval rises and commenting falls) — different units (project coverage vs individual reviewer behavior), possibly both true: projects institute review gates while individual reviewers inside them rubber-stamp more (Review as the Control Point).
  3. The sign flips under definitional choices. Whether agent PRs are reviewed more or less independently than human PRs depends on whether the invoking developer counts as an independent reviewer (agent-as-author) or as the author self-reviewing (agent-as-tool) — CMU's own warning that surface telemetry, vendor or not, cannot adjudicate without a causal model (Telemetry vs. Survey Measurement).
  4. The population gap is unbridged. There is still no non-vendor enterprise telemetry; open-source convergence may not transfer to enterprises, and Faros's enterprise deterioration may not transfer to open source. Neither dataset can falsify the other where it lives.

Verdict#

Read as a contradiction, the divergence dissolves: a delta in unreviewed-PR count across adoption depth (enterprise, all PRs, vendor) and a falling unreviewed share over calendar time (open source, agent PRs, non-vendor) are answers to different questions. The coherent joint reading: AI-driven volume growth raises the absolute amount of under-reviewed code even as teams — first the highest-volume ones, where the pressure surfaced first — learn to triage review by risk, pulling the agent no-review rate down to the human baseline. What survives as a real disagreement is prognostic, not empirical: Faros's fixed-sign "foundations are buckling and volume will break the gates" vs CMU's "the team sets the sign through process adaptation." The next observable test is whether the ~14% convergence holds as agentic authoring's PR share climbs past single digits — a trend line worth re-checking against any post-2026 re-scrape or future Faros/DORA edition.

Sources#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 4
  • Acceleration Whiplash

    Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…

  • Open Questions Backlog

    _396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…

  • Review as the Control Point

    Agarwal et al. (CMU, arXiv 2607.07980): a 26-construct/67-relationship causal theory synthesized from 3,100 coded pract…

  • Telemetry vs. Survey Measurement

    Faros 2026: perception lags reality, so survey-based engineering research (DORA) misses downstream AI damage that syste…

Related articles
  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Acceleration Whiplash

    Faros 2026: AI floods a human-paced SDLC with output it can't absorb — throughput up (tasks +34%, epics +66%), quality…

  • Agentic Coding Work-Composition Shift

    Anthropic's 400K-session telemetry, Oct 2025→Apr 2026: as models improved, the share of sessions fixing broken code fel…

  • Agentic Technical Debt

    Debt that *compounds* (not just accumulates) because each agentic-coding session re-derives architectural decisions wit…

  • Review as the Control Point

    Agarwal et al. (CMU, arXiv 2607.07980): a 26-construct/67-relationship causal theory synthesized from 3,100 coded pract…