Learn in Public
Reading
The work in these essays stands on a lot of other people's. This is the running library underneath it - the papers and references the methodology is built on, with a line on why each one matters. It is the same bibliography the essays cite, kept in one place so you can read past my framing to the sources themselves.
Evaluation & LLM-as-a-judge
How to score model output reliably - panels over single judges, the reliability of LLM judges, and where consensus helps and where it stops.
-
Verga, P. et al. (2024). “Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.” arXiv:2404.18796.
A panel of smaller, diverse models (PoLL) outperforms a single large judge at roughly 7x lower cost and shows less intra-model bias, because disjoint model families do not share the same blind spots.
-
“An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability.” (2025). arXiv:2506.13639.
Mean-of-scores tracks human judgment better than median or majority voting; sampled decoding beats greedy; extreme-anchor rubrics are nearly as good as full ones. A practical map of which judge-design choices actually move reliability.
-
“Self-Consistency Is Losing Its Edge: Diminishing Returns and Rising Costs in Modern LLMs.” (2025). arXiv:2511.00751.
On strong models, self-consistency gains are small (about 0.4-1.6%) and plateau by 10-15 samples while cost scales linearly - consensus tightens the spread, not the center.
-
Haldar, R. & Hockenmaier, J. (2025). “Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks.” arXiv:2510.27106 (EMNLP 2025).
LLM judges have low intra-rater reliability across runs - "almost arbitrary in the worst case" - which is why judge variance has to be measured, not assumed away.
-
Zheng, L. et al. (2023). “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.” arXiv:2306.05685 (NeurIPS 2023 Datasets and Benchmarks).
The canonical LLM-as-judge study - names position, verbosity, and self-enhancement bias as systematic dispositions judges carry, not prompt artifacts, and shows a strong judge still matches human preference ~80% of the time despite them.
-
Wang, L. et al. (2026). “Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment.” arXiv:2602.05110.
In a rubric-scoring task, LLM judges over-score risk by about +0.46 against a panel of human experts, and the lean is model-dependent rather than uniform - independent replication of a directional over-score of roughly half a bin.
-
Kohli, G. (2026). “Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels.” arXiv:2605.29800.
A panel of nine frontier judges from seven model families supplies only about two independent votes' worth of information, because models err on the same items - why a jury cannot average away a prior its members share.
-
Perez, J. et al. (2024). “When LLMs Play the Telephone Game: Cultural Attractors as Conceptual Tools to Evaluate LLMs in Multi-turn Settings.” arXiv:2407.04503 (ICLR 2025).
In iterated transmission chains, small per-step biases amplify and pull text toward stable attractor states - why repeated re-derivation drifts toward the generic rather than wandering randomly.
-
Mohamed, A., Geng, M., Vazirgiannis, M. & Shang, G. (2025). “LLM as a Broken Telephone: Iterative Generation Distorts Information.” arXiv:2502.20258 (ACL 2025).
When a model repeatedly processes its own output, distortion accumulates and compounds with chain length - direct evidence that each hop of a relay degrades the signal.
-
Ma, W. et al. (2023). “(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs.” arXiv:2311.11123.
The base rate that makes eval gating non-optional: 58.8% of prompt+model combinations lost accuracy across API updates, and 55% of updates helped some prompts while hurting others on the same task - there is no uniform "better."
-
Liang, L. et al. (2025). “What Prompts Don't Say: Understanding and Managing Underspecified Prompts.” arXiv:2505.13360.
Underspecified prompt requirements regressed across model updates at roughly twice the rate of explicit ones - the quantified case that a tight contract in the prompt is migration insurance.
Measurement & validity
The older science underneath the scores: comparative judgment, paired-comparison models, construct validity, and when a number is even allowed to be averaged.
-
Krippendorff, K. “Content Analysis: An Introduction to Its Methodology (and Krippendorff's alpha).”
A reliability coefficient where alpha >= 0.80 is "reliable enough to draw conclusions" and around 0.667 supports only tentative ones - a principled gate for when agreement is strong enough to trust.
-
MeasuringU; Statistics By Jim “Can You Take the Mean of Ordinal Data? / Analyzing Likert Scale Data.”
Averaging ordinal codes assumes equal intervals that do not exist; median and quantile summaries are preferred. Why a Low/Med/High scale should not be averaged into a mean.
-
Thurstone, L. L. (1927). “A Law of Comparative Judgment.” Psychological Review 34(4):273-286.
A latent trait can be recovered from comparisons, which absolute ratings compress (scale-usage bias) - the theoretical basis for preferring comparative or anchored judgments over absolute bins.
-
Bradley, R. A. & Terry, M. E. (1952). “Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons.” Biometrika 39(3/4):324-345.
The paired-comparison model that recovers a latent scale from pairwise wins - the practical tool for turning comparisons back into a continuous score.
-
Messick, S. (1995). “Validity of Psychological Assessment.” American Psychologist 50(9):741-749.
Construct under-representation and construct-irrelevant variance as the core threats to measurement validity - the frame for "your instrument may be blind to part of what it is scoring."
-
Morandi, A. (2026). “Two Ways to De-Bias an LLM-as-a-Judge: A Continuous-Score Comparison of Hierarchical Bayesian Calibration and Neural-ODE Score Transport.” arXiv:2605.09227.
Post-hoc calibration leaves the cheap judge in place and fits a transformation from its raw scores toward a human-anchored target - the recalibrate-don't-reprompt move, formalized.
-
Colaco, A. G. & Lahjouji, N. (2026). “What to Keep, What to Forget: A Rate-Distortion View of Memory Compaction in LLMs and Agents.” arXiv:2607.08032.
Irreversible lossy compaction cannot re-derive dropped details, and each event composes its loss with the last, so end-task error grows super-linearly in the number of compaction steps - the multiplicative-compression mechanism, stated formally.
-
Łajewska, W. et al. (2025). “Understanding and Improving Information Preservation in Prompt Compression for LLMs.” arXiv:2503.19114.
Compression methods routinely fail to preserve key details, and the loss lands hardest on multi-hop tasks requiring aggregation - condensing hops shed exactly the information downstream stages most need.
-
Lipsitch, M., Tchetgen Tchetgen, E. & Cohen, T. (2010). “Negative Controls: A Tool for Detecting Confounding and Bias in Observational Studies.” Epidemiology 21(3):383-388.
A negative control is a condition whose result you can predict in advance if your hypothesis holds - a wrong-signed outcome falsifies it. The epidemiologist's formalization of "run the experiment whose answer you already know."
-
Cronbach, L. J. & Meehl, P. E. (1955). “Construct Validity in Psychological Tests.” Psychological Bulletin 52(4):281-302.
The paper that defined construct validity and separated it from reliability: a measure can be perfectly consistent and still not measure the construct you named - reliability is not validity.
Determinism & long context
Why you cannot simply dial in reproducibility, and how long contexts degrade in ways a bigger window does not fix.
-
OpenAI Developer Community “Temperature in GPT-5 models.”
Reasoning models reject a temperature knob - only the default is accepted - so you cannot buy determinism by turning temperature down.
-
Microsoft Learn “How to generate reproducible output with Azure OpenAI.”
A seed is best-effort: identical output is not guaranteed across system or hardware changes. Reproducibility has to be engineered at the seams, not toggled on.
-
Liu, N. F. et al. (2023). “Lost in the Middle: How Language Models Use Long Contexts.” arXiv:2307.03172 (TACL 2024).
Models use the beginning and end of a long context well but degrade sharply on information buried in the middle - so a longer context window is not the same as a usable one.
-
Modarressi, A., Deilamsalehy, H., Dernoncourt, F. et al. (2025). “NoLiMa: Long-Context Evaluation Beyond Literal Matching.” arXiv:2502.05167 (ICML 2025).
When needle and question share no literal overlap, 11 of 13 models fall below half their short-context accuracy by 32K tokens (GPT-4o: 99.3% to 69.7%) - degradation well before the window is full, so more context is not more usable context.
-
Chen, L., Zaharia, M. & Zou, J. (2023). “How is ChatGPT's Behavior Changing Over Time?.” arXiv:2307.09009, published in HDSR.
GPT-4's prime-identification accuracy fell 97.6% to 2.4% between the March and June 2023 versions - same endpoint, same model name, silently different behavior. The founding receipt for continuous monitoring of hosted models.
Retrieval & context
What actually helps a model answer from your data - and where retrieval pipelines quietly lose the signal they were built to carry.
-
Lewis, P. et al. (2020). “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” arXiv:2005.11401 (NeurIPS 2020).
The founding RAG paper: a parametric model paired with a non-parametric index for open-domain question answering - retrieval earns its keep precisely when you cannot point in advance to the passage that answers the question.
-
RAGFlow (2025). “From RAG to Context.”
A retrospective from inside the RAG camp conceding that for fixed, well-structured data, much simpler approaches than retrieval suffice - the tool-fit framing stated by the tool's own advocates.
-
AI21 Labs (2025). “RAG and structured data.”
A retrieval vendor's own admission that chunk-based semantic retrieval is suboptimal and ineffective at scale for tabular and structured data.
-
“RAG "Hype" vs. Reality.” (2025).
A practitioner catalog of RAG's operational limits that keeps returning to the same bottleneck: the whole system is capped by retrieval quality, and chunking routinely shreds document structure to hit a token budget.
-
Shi, F. et al. (2023). “Large Language Models Can Be Easily Distracted by Irrelevant Context.” arXiv:2302.00093 (ICML 2023).
Adding irrelevant context to a problem the model otherwise solves sharply degrades accuracy - direct evidence that extra input is not neutral and can actively pull a model off, not just fail to help.
Tooling & the declarative turn
The frameworks and build systems the essays lean on - from Make and the System R optimizer through dbt and Dagster to typed agent contracts. Declare the what; the engine derives the how.
-
Google Cloud “Document AI: full processor and detail list.”
Every stage of the intelligent-document-processing spine ships as a discrete, named cloud processor - classifier, splitter, parser, custom extractor - evidence that extract-to-schema is settled industry vocabulary, not a bespoke invention.
-
OpenAI “Structured Outputs.”
Constrained decoding compiles a JSON schema into a grammar and restricts token choices at decode time, so output conforms to the schema by construction - extraction stopped being a hope and became a guarantee about structure.
-
Liu, J. “Instructor: structured outputs for LLMs.”
The library that popularized schema-first LLM extraction - typed, validated outputs as the default interface to a model rather than free text.
-
Anthropic (2024). “Introducing the Model Context Protocol.”
The protocol announcement: MCP collapses the N-by-M integration explosion toward N plus M by giving every client and every tool one shared protocol.
-
Model Context Protocol (2025). “MCP Specification (revision 2025-11-25): Transports.”
The current spec revision: stdio and Streamable HTTP as the supported transports, with HTTP+SSE deprecated and retained only for backwards compatibility.
-
Anthropic (2025). “Code execution with MCP.”
The protocol author's own numbers: one workflow drops from 150,000 tokens to 2,000 - a 98.7% cut - by not loading every tool definition into context and letting the model call a small generated API instead.
-
Anthropic (2025). “Agent Skills.”
The lighter primitive: SKILL.md instruction files loaded by progressive disclosure, leaning on the filesystem and shell the model already knows rather than injecting a wall of tool schemas.
-
Ronacher, A. (2025). “Skills vs Dynamic MCP Loadouts.”
A well-designed Sentry MCP server consumes roughly 8,000 tokens of tool definitions loaded up front, and definitions get trimmed and rewritten between versions - the API-stability problem stated from inside the ecosystem.
-
Ngiam, J. “MCPs, CLIs, and skills.”
The clearest statement of where MCP uniquely earns its keep: non-developers invoking tools from inside a chat client, where there is no shell in the loop.
-
Willison, S. (2025). “The Lethal Trifecta for AI Agents.”
Private data, untrusted content, and the ability to communicate externally - the combination with no reliable patch, whose only mitigation is not assembling it in the first place.
-
Taraghi, M., Morovati, M. M. & Khomh, F. (2026). “Real Faults in Model Context Protocol (MCP) Software: A Comprehensive Taxonomy.” arXiv:2603.05637.
The first large-scale taxonomy of MCP-server faults - five high-level categories, validated by practitioner survey - empirical evidence that the layers you operate are layers that fail.
-
Microsoft (2026). “Microsoft Agent Factory.” Microsoft AI.
The hybrid market structure in one primary source: managed agent platforms are sold alongside training, forward-deployed engineering support, and a partner marketplace. Productizing the factory can create a delivery ecosystem rather than remove one.
-
Google Cloud (2025). “Vertex AI Agent Builder overview.” Google Cloud documentation.
A representative full-lifecycle factory: samples and tools, an agent development kit, managed deployment and scaling, evaluation, identities, and security controls are becoming platform capabilities rather than bespoke plumbing.
-
Feldman, S. I. (1979). “Make - A Program for Maintaining Computer Programs.” Software - Practice and Experience 9(4).
Targets and prerequisites as declared artifacts; build order - and what NOT to rebuild - derived. Incrementality and parallelism fall out of the DAG.
-
Selinger, P. G. et al. (1979). “Access Path Selection in a Relational Database Management System.” SIGMOD 1979.
The System R cost-based optimizer - the canonical deterministic derivation engine, and the ancestor of every EXPLAIN plan ever debugged.
-
Miller, P. (1997). “Recursive Make Considered Harmful.” AUUG 1997.
The classic failure mode of a declarative engine given incomplete declarations: partitioned Makefiles hand the engine a fractured DAG, so it derives wrong or slow builds. The fix is total knowledge, foreshadowing Nix and Bazel.
-
Harris, R. (2018). “Virtual DOM is pure overhead.” Svelte blog.
The honest counterweight to the React story: the derivation step itself has a cost the imperative version never paid. Svelte compiles the declaration to imperative code ahead of time instead.
-
Handy, T. et al. (2016). “The dbt Viewpoint.” dbt documentation.
The founding philosophy of models-as-declarations: analytics code should be version-controlled, tested, and declarative-first; the DAG is derived from ref() calls.
-
Dagster Labs “What Is a Software-Defined Asset.” Dagster glossary.
The canonical definition: "a description, in code, of an asset that should exist and how to produce and update it."
-
Schrock, N. (2022). “Rebundling the Data Platform.” Dagster blog.
The launch argument for software-defined assets: orient orchestration around assets rather than tasks, and a single surface of lineage, observability, and quality monitoring falls out.
-
Ryza, S. (2022). “Declarative Scheduling for Data Assets.” Dagster blog.
"You haven't scheduled any jobs, Airflow DAGs, or Prefect Flows. You've just declared how your data flows and when you expect it to be up-to-date." Scheduling derived from declared freshness.
-
Khattab, O. et al. (2023). “DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.” arXiv:2310.03714.
The closest intellectual ancestor of declarative agents: typed signatures declared, prompts compiled and optimized against metrics. "Programming - not prompting - language models."
-
Wang, X. et al. (2024). “Executable Code Actions Elicit Better LLM Agents.” ICML 2024, arXiv:2402.01030.
The code-as-action precedent: replacing JSON tool calls with executable Python as the unified action space, up to ~20% higher success across 17 models. The execution model NOOA builds on.
-
Furgale, P. et al. (NVIDIA) (2026). “NVIDIA-labs OO Agents: Native Python Object-Oriented Agents.” arXiv:2607.20709.
The case study: an agent is a Python object - docstrings are prompts, type annotations are contracts, and methods with `...` bodies are implemented by the model at runtime. Code at github.com/NVIDIA-NeMo/labs-OO-Agents.
-
OWASP CycloneDX (2023). “CycloneDX v1.5: ML-BOM support.” CycloneDX specification.
The SBOM standard grew model and dataset component types plus a modelCard object in June 2023 - the inventory half of the model-BOM problem is standardized; the sourcing-continuity half is not.
-
Linux Foundation Research (2024). “Implementing AI Bill of Materials (AI BOM) with SPDX 3.0.” LF Research implementation guide.
SPDX 3.0's AI and Dataset profiles applied: model type, training information, limitations, energy use, safety assessments. Adoption is early - generators and prototypes, not yet procurement-grade demand.
-
Backstage Project Authors “System Model.” Backstage software catalog documentation.
Shows how machine-readable components, APIs, resources, and relationships can support discovery and dependency visibility; an agent estate would extend this pattern with model routes, evidence, permissions, and operating ownership.
Software history & abstraction
The older papers and essays on how abstractions win, leak, and come back around - the fifty-year context the current agent moment rhymes with.
-
“The Six Sigma Agent: Achieving Enterprise-Grade Reliability in LLM Systems Through Consensus-Driven Decomposed Execution.” (2026). arXiv:2601.22290.
Maps Six Sigma onto agents as redundancy and voting - a clean derivation that is load-bearing on an independence assumption LLM samples do not satisfy.
-
Kephart, J. O. & Chess, D. M. (2003). “The Vision of Autonomic Computing.” IEEE Computer 36(1):41-50.
Named the shape twenty years early: an autonomic manager running Monitor-Analyze-Plan-Execute over shared Knowledge - the agent loop as closed-loop feedback control.
-
NIST/SEMATECH “e-Handbook of Statistical Methods: Process or Product Monitoring and Control.”
The reference text for statistical process control - control limits, process capability, and what it takes to run an imperfect, variable process to a spec.
-
“Statistical process control for drifting ML systems.” (2024).
Prior art for applying SPC-style monitoring to drifting ML systems - it helps by detecting the shift, not by pretending the process is stationary.
-
Google SRE “Embracing Risk (Site Reliability Engineering, ch. 3).”
The error-budget move: decide up front how much unreliability you can spend, measure consumption, and stop spending when the budget is gone - tolerances as a design target.
-
Goldratt, E. M. & Cox, J. (1984). “The Goal: A Process of Ongoing Improvement.” North River Press.
An hour saved at a non-bottleneck is a mirage; an hour gained at the bottleneck is gained for the whole system - the constraint frame the prompt-polishing habit violates.
-
Goldratt, E. M. (1990). “Theory of Constraints.” North River Press.
The five focusing steps - identify, exploit, subordinate, elevate, repeat - and the warning that the constraint moves after you elevate it, so you have to go find it again.
-
Kim, G., Humble, J., Debois, P., Willis, J. & Forsgren, N. (2021). “The DevOps Handbook (2nd ed.).” IT Revolution.
Traces flow, feedback, and continual learning back through Lean and the Theory of Constraints to Deming - one lineage about running an imperfect, variable pipeline fast and safely.
-
Kim, G. (2012). “The Three Ways: The Principles Underpinning DevOps.” IT Revolution blog.
Flow, feedback, continual experimentation - the Three Ways fall onto the agent loop's own arrows almost without translation.
-
Reinertsen, D. G. (2009). “The Principles of Product Development Flow: Second Generation Lean Product Development.” Celeritas Publishing.
The economics of queues, batch size, and variability in development work - why single-piece flow with fast feedback beats big batches, in agent loops as in factories.
-
Cemri, M. et al. (2025). “Why Do Multi-Agent LLM Systems Fail?.” arXiv:2503.13657.
The first empirically grounded failure taxonomy (MAST) - hundreds of hand-annotated traces across seven multi-agent frameworks show failures concentrate in system design, inter-agent misalignment, and task verification rather than in raw model capability.
-
Yuan, D. et al. (2014). “Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-Intensive Systems.” OSDI '14 (USENIX).
Across 198 production failures in five distributed systems, 92% of the catastrophic ones traced to incorrect handling of non-fatal errors the software had already caught, about a third to trivial mistakes like an empty catch block - the empirical spine for "a swallowed error is the dominant failure mode."
-
Candea, G. & Fox, A. (2003). “Crash-Only Software.” HotOS IX (USENIX).
Argues for components whose only stop is a crash and only start is recovery - a clean crash-and-recover is safer than a component limping in an unknown state. The systems-tradition case for "make failure loud" over silent fallback.
-
Dixit, H. D. et al. (2021). “Silent Data Corruptions at Scale.” arXiv:2102.11245 (Meta).
Hardware faults that no error-reporting mechanism flags propagate up the stack and surface as application-level problems months later - the one-layer-down analogue of silent success, at fleet scale.
-
Xie, Y. et al. (2026). “From Spark to Fire: Modeling and Mitigating Error Cascades in LLM-Based Multi-Agent Collaboration.” arXiv:2603.04474.
Minor inaccuracies propagate through a multi-agent dependency graph and solidify into system-level false consensus - the same crush, generalized from a chain to a graph.
-
Bronson, N., Aghayev, A., Charapko, A. & Zhu, T. (2021). “Metastable Failures in Distributed Systems.” HotOS '21 (ACM).
Names the retry-cascade shape precisely: a triggering event pushes a system into an overloaded state that a sustaining feedback loop (often retries) keeps alive, so it stays down after the trigger clears and timeouts or circuit breakers alone do not lift it.
-
Google SRE “Addressing Cascading Failures (Site Reliability Engineering, ch. 22).”
The retry-amplification rule from production: retrying at multiple layers turns one user action into 4^3 attempts, so retry at a single level, use randomized exponential backoff, and cap even that with a per-process retry budget.
-
Huang, R. et al. (2017). “Gray Failure: The Achilles' Heel of Cloud-Scale Systems.” HotOS '17 (ACM).
Introduces differential observability: a fault one component perceives as healthy while another experiences it as failing, so the failure surfaces through the wrong layer and the system's own detectors miss it - why a diagnosis can point confidently at the innocent component.
-
Davenport, T. H. (1998). “Putting the Enterprise into the Enterprise System.” Harvard Business Review.
The durable warning behind the ERP comparison: packaged software integrates the enterprise by carrying assumptions about how the enterprise should work, so implementation is an operating-model decision rather than a software installation.
-
Lacity, M. C. & Willcocks, L. P. (2016). “A New Approach to Automating Services.” MIT Sloan Management Review.
Early RPA field research found that the tool alone was not the operating model: value depended on process selection, executive support, reusable capability, and process experts verifying what the automation actually did.
-
Codd, E. F. (1970). “A Relational Model of Data for Large Shared Data Banks.” Communications of the ACM 13(6).
The founding document of the declarative flip: "Future users of large data banks must be protected from having to know how the data is organized in the machine." Declare relations and a predicate; the system owns access paths.
-
Bachman, C. W. (1973). “The Programmer as Navigator.” ACM Turing Award lecture, CACM 16(11).
The imperative era named by its own champion: the programmer proudly hand-steering through linked records. The perfect artifact of the position the relational model overturned.
-
Kowalski, R. (1979). “Algorithm = Logic + Control.” Communications of the ACM 22(7).
The cleanest theoretical statement of the whole pattern in four words: declare the logic, let the engine supply control.
-
Spolsky, J. (2002). “The Law of Leaky Abstractions.” Joel on Software.
"All non-trivial abstractions, to some degree, are leaky" - and his lead example is SQL: logically equivalent declarative queries whose performance differs by orders of magnitude depending on the optimizer.
-
Stonebraker, M. & Hellerstein, J. (2005). “What Goes Around Comes Around.” Readings in Database Systems.
35 years of data models as a pendulum: every navigational, steps-first revival eventually collapsed back into the declarative model once the engines got good enough.
-
Chishti, Oyinloye & Li (NTNU) (2026). “Test Before You Deploy: Governing Updates in the LLM Supply Chain.” arXiv:2604.27789.
Frames hosted-model updates as a supply-chain governance problem: providers update weights, safety policies, and serving infrastructure "without changing the API endpoints" - behavior changes with no version change, which breaks the core assumption of dependency management.
-
Brynjolfsson, E. & Hitt, L. M. (2000). “Beyond Computation: Information Technology, Organizational Transformation and Business Performance.” Journal of Economic Perspectives 14(4).
A review of firm-level evidence showing that organizational complements materially shape the value realized from information-technology investment; it is historical lineage, not direct evidence about agents.
-
Hammer, M. (1990). “Reengineering Work: Don't Automate, Obliterate.” Harvard Business Review.
The canonical argument for redesigning obsolete work before using technology to accelerate it; cited as intellectual lineage rather than a universal prescription.
Industry & practice
Announcements, reporting, and field notes that mark where the industry is actually moving - the receipts behind the commercial claims.
-
Yao, S. et al. (2022). “ReAct: Synergizing Reasoning and Acting in Language Models.” arXiv:2210.03629.
The paper that wrote down the agentic loop - reason, act, observe - as an explicit structure rather than an emergent behavior.
-
Shinn, N. et al. (2023). “Reflexion: Language Agents with Verbal Reinforcement Learning.” arXiv:2303.11366 (NeurIPS 2023).
Extends the loop with reflect-and-retry - and quietly presupposes the thing most teams skip: a signal that says the last attempt was wrong.
-
Anthropic (2024). “Building Effective Agents.”
Distinguishes predefined workflows from agents that dynamically direct their own process and tools, and recommends adding agentic complexity only when simpler approaches fall short.
-
National Institute of Standards and Technology (2023). “AI Risk Management Framework Core.” NIST AI Resource Center.
Defines outcomes for targeted application scope, operator proficiency, human oversight, independent review, monitoring, appeal and override, deactivation, recovery, and change management.
-
Data Science Dojo (2026). “Loop Engineering.”
Representative of the 2026 "loop engineering" genre: the right slogan - design the loop, not the prompt - that resolves into a guardrails checklist and stops where the engineering starts.
-
Shaw, S. D. & Nave, G. (2026). “Thinking—Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender.” Wharton working paper; SSRN 6097646; PsyArXiv DOI 10.31234/osf.io/yk25n_v1.
Three preregistered experiments (n=1,372, ~10,000 trials): a confidently wrong AI drags answers 15 points below baseline while raising user confidence 11.7 points, and 73% of wrong answers were accepted without any attempt to override - the paper coins "cognitive surrender."
-
Kosmyna, N. et al. (2025). “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task.” arXiv:2506.08872 (MIT Media Lab).
EEG study: AI-assisted writers showed the weakest, least-distributed neural connectivity of three groups and weaker recall of their own text - suggestive of a mechanism, not proof of permanent decline (small-N, short-horizon, essay writing).
-
Anthropic (2026). “How AI assistance impacts the formation of coding skills.”
Randomized trial, 52 mostly-junior engineers: AI-assisted averaged 50% vs 67% hand-coded on comprehension - but conceptual-inquiry users scored 65%+ while generate-and-delegate users scored under 40%. The mode of use decided the outcome, not the tool.
-
Bainbridge, L. (1983). “Ironies of Automation.” Automatica 19(6):775-779.
The classic irony: automating a process erodes the operator skill needed precisely in the abnormal moments when the human must take over.
-
Parasuraman, R. & Manzey, D. H. (2010). “Complacency and Bias in Human Use of Automation: An Attentional Integration.” Human Factors 52(3):381-410.
Automation complacency and bias appear in experts as readily as in novices and are not cured by simple practice - the out-of-the-loop operator is a known hazard with a known shape.
-
Osmani, A. (2026). “Cognitive surrender and comprehension debt.”
Sharpened "cognitive surrender" for practitioners and named comprehension debt: the widening gap between the volume of code in your name and the amount of it you actually understand.
-
Kahn, J. (2026). “Anthropic's Boris Cherny doesn't write code by hand anymore.” Fortune (Brainstorm Tech, Aspen, June 11, 2026).
The primary source for the Cherny quote: "I haven't written a line of code by hand in, I think, eight months now."
-
Cherny, B. (2026). “"As engineering, product, design, DS, etc. melt into a new kind of role..." [Post].” X (June 28, 2026); cross-posted by the author to Threads the same day.
Primary source for the five Claude Code team archetypes: prototyper, builder, sweeper, grower, maintainer. A hedged, forward-looking observation ("what I think is five archetypes") about one team, not a staffing study. Author fallback copy, publicly readable without an X account: https://www.threads.com/@boris_cherny/post/DaJgVFVj2PB/
-
Zaharia, M. et al. (2024). “The Shift from Models to Compound AI Systems.” Berkeley AI Research (BAIR) blog.
Defines a compound AI system as one that tackles tasks using multiple interacting components - model calls, retrievers, external tools - and argues state-of-the-art results now come from the system, not a monolithic model.
-
Schmidt, J. (2025). “Trading Margin for Moat: Why the Forward Deployed Engineer Is the Hottest Job in Startups.” Andreessen Horowitz.
The wave's own economics, said out loud: trade gross margin for control of the deployment layer, because the implementation-heavy companies (Salesforce, ServiceNow, Workday) started margin-ugly and ended up as systems of record. Also counts 22 of OpenAI's 311 open roles as forward-deployed/solutions engineering.
-
Orosz, G. (2025). “What are Forward Deployed Engineers, and why are they so in demand?.” The Pragmatic Engineer.
The role's documented history: Palantir created the title in the early 2010s and called the people who held it "Deltas" - and until around 2016 the company had more forward-deployed engineers than conventional software engineers.
-
Palantir Technologies Inc. (2026). “Annual Report for the Year Ended December 31, 2025.” Form 10-K, U.S. Securities and Exchange Commission.
Bounds the economic claim: Palantir reports subscription software, O&M, and professional services in one productized operating model, including customer-facing configuration, training, ontology, and data-modeling support. It reported 82% consolidated gross margin in 2025 but does not disclose an FDE-specific P&L.
-
PYMNTS (2026). “Forward-Deployed Engineers Emerge as One of AI's Fastest-Growing Jobs.” PYMNTS, citing the Financial Times.
Secondary report of a Financial Times analysis of Indeed postings: monthly listings for the title reportedly grew more than 800% between January and September 2025. The underlying series is not reproduced, so the essay uses this only as a directional signal of title demand.
-
SAP SE (2015). “SAP Solution Manager 7.1: ALM Processes in Detail.” SAP Help Portal.
Documents ASAP as a repeatable implementation method with roadmaps, customer-specific blueprints, configuration, testing, operations, and reusable implementation content.
-
Infosys Technologies Limited (2003). “Annual Report on Form 20-F for Fiscal 2003.” Infosys investor filing.
A primary-source receipt for services industrialization: Infosys describes decomposing projects across client sites and offshore centers, training rapidly deployable professionals, reusing knowledge, and executing components where they are most cost-effective.
-
Accenture plc (2025). “Annual Report for Fiscal 2025.” Form 10-K, U.S. Securities and Exchange Commission.
Describes standardized processes, methods, tools, automation, global delivery, industry specialization, and cost advantages as inputs to scalable, price-competitive services.
-
Deng, X. (N.) (2010). “Acting as Translators between Consultants and Users in ERP Implementation: An Exploratory Study of Analysts' Boundary Spanning Expertise.” International Research Workshop on IT Project Management.
Direct precedent for the analyst-builder thesis: effective boundary spanners need overlapping business and technical knowledge, plus the standing to probe assumptions and challenge the status quo across both groups.
-
Ko, D.-G., Kirsch, L. J. & King, W. R. (2005). “Antecedents of Knowledge Transfer from Consultants to Clients in Enterprise System Implementations.” MIS Quarterly 29(1).
A matched-pair study across 96 ERP projects. It treats the client's ability to apply implementation knowledge and maintain the system independently as an expected outcome, grounding capability transfer as more than a consulting slogan.
-
OpenAI (2026). “OpenAI launches the OpenAI Deployment Company to help businesses build around intelligence.” OpenAI.
A primary-source receipt that the model provider is rebuilding the implementation layer: DeployCo launched with more than $4 billion of initial investment and approximately 150 forward-deployed engineers and deployment specialists from Tomoro.
-
Ode with Anthropic (2026). “Anthropic, Blackstone, and Hellman & Friedman Introduce Ode with Anthropic, an Enterprise AI Services Firm.” Ode.
A second primary-source receipt that frontier-model access is being paired with a dedicated delivery institution. Ode combines Anthropic engineers with the acquired Fractional AI team and positions itself as an end-to-end partner from roadmap through deployment. The official announcement does not disclose a dollar value.
-
Bellan, R. (2026). “Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models.” TechCrunch.
Reports Ode as a $1.5 billion company. The essay and figure keep this separate from DeployCo's officially disclosed initial investment because company value and invested capital are not directly comparable.
-
OpenAI (2026). “OpenAI API Pricing.” OpenAI.
Primary-source evidence that API usage is metered and billed through input, cached-input, and output tokens. It establishes the recurring-usage mechanism without disclosing DeployCo's economics or proving that token pull-through motivated the investment.
-
Anthropic (2026). “Claude Model Pricing — All Platforms.” Anthropic.
Primary-source rate card showing Claude API pricing per million input, output, and cache tokens. It supports recurring model-consumption economics while leaving Ode's services and referral economics undisclosed.
-
Prefect (2026). “Prefect Acquires Dagster Labs.” prefect.io, July 13, 2026.
The commercial tell: the two leading Airflow successors merge, mapping Dagster as the "outcomes layer," Prefect as the "execution layer," and FastMCP as the "access layer" of an agent-orchestration platform.
-
Schrock, N. (2026). “Prefect is Acquiring Dagster.” Dagster blog, July 13, 2026.
The founder's own note on the merger: Prefect's dynamic workflows are 'newly relevant in the agentic era'; Dagster brings partitioning, lineage, cataloging, and scheduling. Schrock steps down.
-
OpenAI “Model deprecations.” OpenAI developer documentation.
The primary source for shutdown dates and notice floors: 6 months for GA models, 3 for specialized variants, as little as 2 weeks for previews. The July 23, 2026 wave retired 18 snapshots including the deep-research models.
-
Tursio Inc. (2025). “Prompt Migration: Stabilizing GenAI Applications with Evolving LLMs.” arXiv:2507.05573.
The best quantified migration war story: a 100%-passing suite dropped to 97-98% on naive model swap; a systematic testbed and prompt restructuring recovered 100% and cut migration effort from several months to two weeks.
-
Anthropic (2025). “Commitments on model deprecation and preservation.” Anthropic research blog, Nov 4, 2025.
The vendor acknowledging the problem: weights of all publicly released models preserved "for, at minimum, the lifetime of Anthropic as a company," plus post-deployment reports. Preservation is not availability - the 60-day retirement floor coexists with it.
-
Microsoft “Azure OpenAI model deprecations and retirements.” Microsoft Learn.
GA models available a minimum of 12 months from launch with at least 60 days notice before retirement - the same weights on a longer contract. Azure lists o3-deep-research retiring ~5 months after OpenAI shut it off.
-
AWS “Amazon Bedrock model lifecycle.” AWS documentation.
At least 12 months on platform and 6 months of Legacy notice, plus a paid Extended Access phase. Claude 3.7 Sonnet lived two months longer on Bedrock than on Anthropic's own API - distribution contract as sourcing terms.
-
Huckins, G. (2025). “The people who lost their AI companions when GPT-4o was retired.” MIT Technology Review, Aug 15, 2025.
The serious account of the #Keep4o episode. A follow-up peer-reviewed study coded the responses: relational attachment, disappointment, grief - model updates as "significant social events."
-
Cognition (2025). “Rebuilding Devin for Claude Sonnet 4.5: Lessons and Challenges.” Cognition engineering blog.
The counter-evidence to "a good harness makes models swappable": a better model "broke our assumptions about how agents should be architected" and required a harness rebuild - which then paid off at 2x speed.
-
OpenAI (2026). “OpenAI to acquire promptfoo.” openai.com, March 9, 2026.
The vendor whose release cadence forces migrations bought the leading open-source migration-testing tool - model-swap evaluation is now strategic infrastructure, by the acquirer's own admission.
-
Willison, S. (2026). “The new GPT-5.6 family: Luna, Terra, Sol.” simonwillison.net, July 9, 2026.
The replacement model, independently assessed: "definitely very competent," state of the art on some benchmarks - and even expert users publicly puzzling over which effort level to run it at. Capability was never the migration problem.
-
Forsgren, N. et al. (2019). “Accelerate State of DevOps 2019.” DORA research report.
Reports an association between heavyweight external change approval and lower software-delivery performance, and recommends peer review plus automation; the finding is correlational and is not evidence to remove consequential review.
-
DORA (2026). “Platform engineering.” DORA capabilities guide, updated January 12, 2026.
Treats an internal platform as a product and recommends measuring adoption, retention, task success, and delivery outcomes. Its research reports correlations and capability guidance, not causal proof for an AI operating model.
-
DORA (2024). “Accelerate State of DevOps Report 2024.” DORA research report.
In a survey of nearly 3,000 technology professionals, a 25% increase in reported AI adoption was associated with better documentation, code quality, code review speed, and approval speed, but lower delivery throughput and stability. The estimates are correlational and the report charts 89% uncertainty intervals.
-
DORA (2025). “State of AI-assisted Software Development 2025.” DORA research report, version 2025.2.
In a survey of nearly 5,000 technology professionals, the point estimate for AI adoption and software-delivery throughput was positive but its 89% interval narrowly crossed zero; the estimate for delivery instability was clearly positive. The standardized cross-sectional model differs from 2024, so direction can be compared but coefficient magnitude cannot be treated as a longitudinal trend.
-
Model Evaluation & Threat Research (METR) (2026). “Measuring AI Ability to Complete Long Tasks: Time Horizon 1.1.” METR research dashboard, updated May 8, 2026.
Estimates the human-expert task duration at which models complete 228 well-specified technical tasks with 50% or 80% success. The results are not measures of autonomous job duration, carry wide uncertainty intervals, and are unreliable above 16 hours with the current task suite.
-
National Institute of Standards and Technology (2023). “Artificial Intelligence Risk Management Framework (AI RMF 1.0).” NIST AI 100-1.
A voluntary, use-case-agnostic framework covering roles, inventory, deployment-context evaluation, monitoring, incident response, recovery, and decommissioning.
-
Fowler, M. (2024). “Strangler Fig Application.” martinfowler.com.
Explains incremental modernization through outcome definition, decomposition, transitional architecture, and gradual replacement rather than a single high-risk cutover.
-
Google SRE “Canarying Releases.” The Site Reliability Workbook.
Describes staged exposure, control comparison, evaluation, and rollback as a production-release discipline. Agentic work needs additional semantic and operating-consequence evaluation.
These sources underpin the Field Guide - the essays show what they look like in production.