Charts on this page visualize AI exposure scores across 404 energy positions. Each exhibit can be explored interactively. Data is available in the downloadable CSV dataset.
The Energy Decision Stack / September 2026

The people and the money are in different buildings.

AI’s first opening in energy is the work around the decision: the reserve report, the lender package, the regulatory filing. We mapped 404 positions to find where that opening matters most.

404 positions mapped373 people-roles60 BLS / industry anchors344 formulaic employment estimatesOpen data & methods
September field notes / 07 SEP 2026

The opportunity grew.
The constraints got clearer.

Cheaper reasoning meets a slower physical world. The new evidence strengthens the case for better preparation — and makes it harder to claim that better paperwork alone delivers more power.

01 / Grid access
2,061 GW queued

Generation and storage at end-2025, down from roughly 2,300 GW. Completed projects took a median of more than five years to reach operation in 2025. A smaller queue can still move slowly.

Berkeley Lab · May 2026 ↗
02 / Electricity demand
485 → 950 TWh

Global data-center electricity use in 2025 and the IEA’s 2030 projection. This includes all data centers. Demand nearly doubles while power, chips, and grid access constrain delivery.

IEA · April 2026 ↗
03 / Capital decisions
$3.4 trillion

Estimated global energy investment in 2026. The case for AI is strongest where a better analysis can improve a material decision. Realized returns still have to be demonstrated.

IEA · 2026 outlook ↗

What this edition updates: external evidence, model pricing, charts, and pilot economics. The April role scores and employment assumptions are preserved as a dated benchmark. The earlier 60–70% estimate for document work in queue delays is retired: it lacked project-level validation. Read the change log and corrections →

Research summary

The Energy Decision Stack. Raj Mistry, Sunya Research. Published April 2026; September edition dated September 7, 2026.

The original 404-position benchmark contains 373 roles, 24 workflows, and 7 artifacts. About 90% of exposure-weighted wage dollars lie above the physical layer under its assumptions; about 89% in the people-roles-only view and about 74% in the anchor-only view. These are modeled sample results, not measured national workforce shares or statistical confidence bounds.

Employment uses 60 BLS/industry anchor rows and 344 formulaic estimates; compensation uses six layer proxies. Scores are single-analyst estimates and have not been validated against before-and-after operating results.

September evidence: 1,312 GW generation and 749 GW storage queued at end-2025; median time to operation exceeded five years for projects completed in 2025. A separate historical cohort shows 13% of requested capacity reaching operation. The queue excludes data-center load requests. Global data-center electricity demand was 485 TWh in 2025, with an IEA projection near 950 TWh in 2030. Global energy investment is estimated at $3.4T in 2026.

The 60–70% queue-document-time claim is retired. More scenarios may improve decisions, but no causal link from scenario count to investment returns has been established here. The proof page is a synthetic storyboard, not a measured model run. The pilot calculator values net released capacity after review, API, software, and implementation costs; it does not predict layoffs or cash savings.

The thesis

We tried to figure out which energy jobs AI would kill and accidentally discovered the people and the money are in different buildings.

We scored 404 energy positions to see where AI hits hardest, because the public conversation about this is just people yelling "robots!" at each other without a spreadsheet. Turns out: the field is mostly fine. The roughnecks, the linemen, the turbine techs — they score a 3.5 out of 10 on compressibility, which is a fancy way of saying AI is not great at climbing transmission towers. But the people assembling 200-page lender presentations and regulatory filings? 6.5 to 8.2. Ninety percent of modeled exposure-weighted wage dollars sit above the field, in the analytical layers, doing prep work for decisions that control billions. The industry keeps buying AI to save $15M in G&A. The $3B capital allocation decision that G&A services is sitting right there. It's like hiring a personal shopper to optimize your grocery bill while your investment portfolio picks itself.

We went down this rabbit hole because nobody had published a spreadsheet. One side of the AI-and-energy-jobs debate says the robots are coming for the roughnecks. The other side says AI will create millions of green jobs. Both sides are guessing. So we scored 404 positions across six organizational layers — compressibility, wage-weighting, stress tests, the whole thing — to figure out what actually compresses and what doesn't.

The answer is not what anyone expected, including us.

Physical operations — the hard-hat jobs the industry worries about — are roughly 29% of headcount but less than 10% of AI-exposed wage dollars. Field roles score 3.5 out of 10 on compressibility. (It turns out "have a human physically be on a drilling platform" is hard to automate. Who knew.) The roles that do compress are all above the field: analysts building reserve reports, coordinators assembling regulatory exhibits, associates packaging lender decks. These are the people whose compressibility scores run 6.5 to 8.2, and whose entire job is to get information into a shape where someone senior can say yes or no.

Preparation, review, and judgment respond differently to AI. The role may persist while the work inside it shifts: fewer hours rebuilding the same package, more hours resolving exceptions and testing alternatives. The size of that shift is an empirical question. A time-and-error log from one recurring workflow will tell you more than a confident percentage applied to an entire job title.

The useful comparison is the capital decision beside the overhead budget. In an illustrative company with $75M of G&A and $3B of development spending, a 20% overhead reduction is $15M; a 2% improvement in investment outcomes is $60M. The arithmetic explains why decision quality deserves attention. It does not establish that AI delivers the 2% improvement. That is what the pilot has to test.

The real problem is the org chart. If the leverage is in capital allocation, then the CFO needs to own AI deployment, not the IT department two reporting layers away from anyone who touches a development decision. Right now most energy companies hand AI to the CIO, the CIO builds a chatbot, the chatbot summarizes meetings, and everyone writes a LinkedIn post about digital transformation. Meanwhile the analysts are still manually building variance tables for a billion-dollar credit facility. Someone got promoted for the chatbot. The variance tables are still in Excel. That probably changes at some point, but the gap between "models can do this" and "your company actually does this" is where all the money is.

Five things we found (30 seconds)

01

The mismatch is the whole story. About 29% of modeled headcount is in the field. Less than 10% of AI-exposed wage dollars are there. That's a 3:1 mismatch between where the people are and where the money is. In finance, they'd call that a mispricing.

02

Titles may survive while work changes. Drafting and reconciliation are candidates for compression; signoff, exception handling, and judgment remain. The benchmark ranks exposure. It does not measure what fraction of a person’s week disappears.

03

The upside may extend beyond efficiency. Cheaper preparation can make room for more alternatives and better challenges to an investment case. That is a hypothesis to test. More scenarios do not, by themselves, establish a better decision.

04

Three candidates for a measured pilot. Treasury/lender readiness, ownership and title, and regulatory preparation combine recurring work with an identifiable reviewer and budget owner. Start with company-controlled preparation and measure the full cost of accepted output.

05

Energy constrains AI’s expansion. The updated queue data records 2,061 GW of proposed generation and storage at end-2025. Preparing studies and applications more efficiently can help, but equipment, construction, financing, and external approvals still determine delivery. Berkeley Lab, 2026.

You're thinking "just give me the bullets." We know.

Jump to section 25 min total read
01Where the money actually is (it's not the field)4 min

29% of energy workers are in the field. Less than 10% of AI-exposed wage dollars are. The slope chart makes this visceral — the lines cross, and ninety cents of every exposed dollar sits above the field in analysts and packagers.

02404 jobs, one chart — find yours3 min

Every position in energy, mapped by headcount and AI exposure. The biggest boxes are the lightest — that's where most people work, and AI barely touches them. The dark boxes are small but expensive. Your job is probably in here.

03Type your job title. See what's left.3 min

Type any energy role and get its compressibility score, what stays human, and what the AI actually does. Thirty roles with individual breakdowns — the specificity is the point.

04Three workflows you can start next quarter4 min

Borrowing-base assembly, regulatory exhibit prep, and type-curve scenario analysis. Three candidate pilots with repeatable inputs, identifiable reviewers, and measurable preparation costs. Returns depend on quality and implementation.

05What AI won't fix (no matter what the vendor says)2 min

Faster packets don't mean faster decisions. Lender committees still meet when they meet. Bad source data breaks everything. And if nobody trusts the output, the review burden goes up, not down. This section is the cold water.

06Eight company types. Eight different starting points.3 min

Not every company gets the same playbook. Majors, independents, and PE-backed platforms each have different leverage points. This maps which company type benefits most from which deployment strategy.

07Why OpenAI needs your filing cabinet3 min

AI needs compute. Compute needs data centers. Data centers need power. Power needs permitting — and permitting is document work. The most advanced technology on earth is bottlenecked by regulatory filings that AI already knows how to compress.

08How exposed is your team? (calculator)2 min

Plug in your team's headcount by role type and get a weighted exposure score, total exposed wage bill, and above-field percentage. Takes 30 seconds. Share the result with your CFO.

09All 404 rows. Every assumption. Every caveat.4 min

Every assumption, every data source, every limitation. The scoring rubric, BLS employment anchors, compensation proxies, sensitivity analysis, and the full 404-row dataset. If you want to check our work, this is where you do it.

Eighteen months ago, you couldn't feed a 200-page reserve report into a model and get anything useful back. You can now. That's the whole timing argument. The gap between "models can do this" and "your competitor already is" tends to close faster than anyone expects.

All scores are modeled estimates. Every claim is tagged by confidence level. Full methodology below.
The core finding
90%
of modeled exposure-weighted wage dollars sit above the field — in analysts, packagers, and preparers
TL;DR

The industry worries about roughnecks. The money says worry about the people writing the spreadsheets the roughnecks run on.

Exhibit 1 · 404 positions across 6 layers

The people are in the field. The money isn't.

About 29% of modeled headcount is in the field. Less than 10% of AI-exposed wage dollars are. Everyone in the industry kind of knows this. Now there's a spreadsheet.

Left: where the people are. Right: where the exposed dollars concentrate. The lines cross. Ninety cents of every AI-exposed wage dollar sits above the field — seventy-four cents even when restricted to BLS-anchored rows only.

What could break this

Faster packets ≠ faster decisions. A compressed borrowing-base package still waits for the VP's calendar, the lender committee, and the bank's own reserve engineers. Maybe 30% of cycle time is work product. The rest is organizational friction AI doesn't touch.

Bad source data breaks everything. Well logs from the 1970s were hand-transcribed. Conflicting spreadsheets, stale type curves, vintage assumptions. AI on bad data produces bad results faster.

Weakly trusted outputs increase review burden. If the model's work isn't trusted, reviewers spend more time checking than they saved. The net effect can be negative until confidence builds.

How robust is the ~90% figure?

Below: what happens when we stress field-layer headcount and compensation — the two inputs most likely to be underestimated. Non-field employment estimates and compensation proxies are held constant; stressing those would require a broader sensitivity test. The pattern holds across the range tested because the compensation gap between layers dominates the math.

Rows: field employment multiplier (how much larger or smaller the physical layer headcount might actually be). Columns: field compensation adjustment. The base case uses the current model assumptions. Even at the extremes — 2× field headcount at 1.5× field pay — the above-field share still shows the majority of exposed dollars sitting outside physical operations.

Extended sensitivity: above-field parameter stress

What happens when we vary the parameters the above-field layers use? Rows: multiplier on all non-Physical layer bases. Columns: multiplier on all non-Physical compensation proxies. Both directions weaken the result by making above-field layers less dominant.

What drives the ~90% — and what survives stress

Three factors build the above-field concentration. The decomposition below shows how much each contributes (under the score-neutral employment model), and the stress tests show what happens when we deliberately weaken the most attackable assumptions.

Factor decomposition

Employment alone
70.5%
+ compensation proxies
79.4%
+ compressibility scores
90.3%

Even with uniform compensation and uniform compressibility, ~71% of estimated headcount sits above the field (under score-neutral employment). Compensation proxies add ~9 percentage points. Compressibility scores add another ~11. The direction is established by headcount distribution alone — the other factors amplify it.

Stress tests

Scenario
Above field
Baseline — all 404 positions, score-neutral employment
90.3%
Flatten compensation to $80K for all layers
(removes the pay-gap amplifier between Physical at $58K and Governance at $180K)
85.2%
Cap compressibility at 7.0 for all rows
(assumes above-field roles are less compressible than scored)
89.0%
Both: flat compensation + capped scores
(harshest combination — removes both amplifiers)
83.4%
Anchor rows only (60 of 404, BLS/industry-sourced)
(excludes all formulaic employment estimates)
74.2%
Roles only — exclude workflows and artifacts (373 of 404)
(addresses the mixed-ontology objection)
89.4%

The above-field share ranges from 74–96% across all stress tests (parametric Monte Carlo: 75.5–96.1%; anchor-row-only test: 74.2%). Even under the harshest tested specification (BLS-anchored rows only, excluding all formulaic estimates), roughly three of every four exposed wage dollars still sit above the field. The directional finding is not an artifact of the employment model, the compensation proxies, or the scoring method.

What could break this

344 of 404 employment estimates use a deterministic formula (layer base × tier multiplier) rather than observed data. Three tier bands (0.75×, 1.0×, 1.25×) assigned by stable hash provide structural spread, not an economic model. The compensation proxies are layer-level averages, not role-specific wages. If physical-layer headcount is substantially larger than modeled, or if above-field compensation proxies are too high, the ~90% figure compresses — but the stress grid shows it stays above 74% even when physical headcount doubles and physical comp increases 50%. The anchor-row test (74.2%, using 60 BLS/industry-sourced rows only) provides the hardest floor. The direction holds across all specifications. The exact percentage is scaffold-dependent — treat it as an interval (74–96%), not a point estimate.

Compressibility scores
404 / 404
All positions scored
Augmented axes
304 / 404
Decision criticality, reasoning demand, company control
Employment anchoring
60 / 404
BLS-sourced · 344 use formulaic estimates

What this does not claim

(a.k.a. the section most research puts in an appendix nobody reads)

Modeled, not observed. Every score is informed estimation. No before/after workflow data exists yet. Validation against deployed systems is planned for a future edition.

Exposure is not replacement. A high compressibility score means the prep stack compresses. The title, the signoff, and the judgment remain human.

Scores rank directionally. A 0.4-point difference between two roles is noise, not signal. Read tiers, not decimals.

Employment and comp are proxies. 60 of 404 positions use BLS/industry employment anchors. The remaining 344 use a formulaic estimate (layer base × tier multiplier) — structural scaffolding, not survey data. Compensation uses six layer-level proxies ($58K–$180K), not role-specific wages. The layer-level pattern is the claim; individual role numbers are not precise.

External approvals stay external. Compressing a borrowing-base packet does not make the lender approve faster. Compressing a rate-case filing does not make the PUC rule sooner. Company-controlled loops compress. Externally governed outcomes do not.

Energy is finance in 1975.

$3.4 trillion in estimated 2026 investment, almost all of it driven by individual judgment, individual spreadsheets, individual analysts running four scenarios when they should run forty.

(The firms doing this in finance in 1975 — Kidder Peabody, Drexel, Salomon Brothers — don't exist anymore.)

Capital allocation gap
40:1
Capital decisions outweigh G&A savings — but AI budgets target G&A
TL;DR

You're optimizing the grocery bill while the investment portfolio picks itself. The leverage is in decisions, not overhead.

Exhibit 2 · 404 positions

404 positions. Same story every time.

Area is headcount. Color is AI exposure. The biggest boxes are the lightest — that's where most people work, and AI barely touches them. The dark boxes are small but expensive. Every strategy deck worries about the big light boxes. The money is in the small dark ones. (It's always in the small dark ones.)

Same mismatch. Every role. Rank by score: one answer. Rank by dollars: completely different list.

Exhibit 3 · Wage-bill ranking

Score tells you what's compressible. Dollars tell you where to start.

So now we know where the money is. But who, specifically? Trade-exception processing and non-op accounting top the scorecard. But they're not where the dollars are. Landman, lineman, petroleum engineer, geologist, production engineer — that's where the money is. The punchline: what's most compressible and what's most expensive are different lists.

The 20 roles where AI exposure meets real money. Longer line = bigger wage pool. Darker circle = more compressible. Notice what's not on this list.

Exhibit 4 · Scenario analysis (hypothesis)

The system surfaces the decision. The engineer applies judgment.

OK so the ranking is clear. What happens when this stuff actually compresses? Two versions of this story. The boring one: AI makes the same analysts faster. The interesting one: the system runs forty scenarios, flags the three that actually change the decision, and delivers them before the cycle starts. Same data, completely different output. The bottleneck moves from "can we assemble the evidence" to "what do we do with it." That second question is worth more. It's also harder. Nobody ever got fired for assembling evidence slowly.

Karpathy's framing applies here directly: "To get the most out of the tools that have become available now, you have to remove yourself as the bottleneck." In energy, the bottleneck today is evidence assembly. When that compresses, the constraint shifts to judgment — which is where you want engineers spending their time.

Two paths (hypothesis — each requires validation)

Path A · Copilot

Same workflow, faster. Analysts run 4 scenarios in 3 days instead of 3 weeks. The decision stays the same. Cost savings: 10–20% on analytical labor.

Path B · Intelligence

Different structure entirely. The system composes 40 scenarios from capabilities, flags the 3 that change the decision, and delivers them before the cycle starts. The decision changes. Value: 10–100× the labor savings.

The gap in this benchmark: it does not contain an observed comparison of these two paths. More scenarios are useful only when they are valid, explore a material uncertainty, and change the decision for a defensible reason. A future operating study should record which alternative changed the recommendation, what error checks passed, and whether the eventual outcome improved.

Annual model-token scenarios per operating unit. Hypothetical — not observed. But the shape of the curve is the argument: demand doesn't flatten when analysis gets cheap. It compounds.

The Jevons trap — and why G&A is the wrong denominator

Here's what most energy companies do first with AI: email drafts. Meeting summaries. Formatting. Which makes sense — nobody gets fired when a model hallucinates a meeting recap. But it misses the point by about 40:1.

A mid-size E&P spends $50–100M on G&A. It spends $2–6B on development capital. That's a 30-to-1 ratio. The entire AI-for-energy pitch is aimed at the smaller number. It's like hiring a personal shopper to optimize your grocery bill while your investment portfolio picks itself.

The Jevons analogy is a hypothesis about demand: as analysis becomes cheaper, firms may choose to ask more questions. The workflow scenarios below illustrate that possibility. They do not demonstrate the size of the response or its effect on investment returns.

Which is why this is not an efficiency story. Saving 20% of G&A on $75M is $15M. A 2% improvement in capital allocation on a $3B development program is $60M — four times the G&A savings the industry's writing press releases about. The industry is optimizing the small number and ignoring the big one.

In 2019, Concho Resources drilled 23 wells at 230-foot spacing on the Dominator pad — roughly a third of what the rest of the Delaware Basin was running. The thesis was more wells per section, more resource recovery. What happened was well interference: production was tracking 38% below the rest of Concho's Wolfcamp projects — and headed toward 45% below pre-drill estimates. The stock dropped 22% in a single day — $4.4 billion in market cap. The team that selected the spacing — reservoir engineers, geologists, a VP, a CEO — earned a combined $3–5M a year. They destroyed a thousand times their own compensation in one decision with inadequate scenario coverage.
Concho Resources, Delaware Basin, 2019

That's the leverage structure of this industry. A handful of knowledge workers control a capital budget that dwarfs their salaries by orders of magnitude. Dominator illustrates the cost of a development decision going wrong. This report does not establish that insufficient scenario coverage caused the result or that AI would have prevented it. Make analysis cheap, and the question changes from "can we afford to run more cases?" to "can we afford not to?"

The same allocation problem appears across the AI supply chain. The IEA estimates that capital expenditure by five large technology companies could rise roughly 75% from more than $400B in 2025. This is total company capex, not a clean measure of AI-only investment. The analytical task is to identify which sites, contracts, and infrastructure choices turn spending into productive capacity. IEA, April 2026.

Illustrative mid-size E&P: $75M G&A and $3B development capital. The 20% overhead and 1–5% allocation changes are sensitivities, not observed or predicted AI returns. Cross-industry examples retain April assumptions. The historical Dominator loss is not included as an AI-avoidable benefit.

The ownership problem

If the value is in the capital allocation decision — not the meeting summary, not the email draft, not the reformatted slide deck — then the person who owns that decision has to own the AI deployment. Not "approve it." Not "sponsor it." Own it. Be the user. Personally.

Right now, most energy companies hand AI to the IT department or to a "digital innovation" team two reporting layers removed from anyone who touches a capital decision. That's how you get chatbots for the help desk and summarizers for the weekly all-hands. Useful. Worth maybe $15M a year. And completely disconnected from the $3B development program where the real leverage sits. Everyone writes a LinkedIn post about digital transformation. Meanwhile the analysts are still manually building variance tables for a billion-dollar credit facility.

The CFO who reviews four scenarios that someone else built is a different person than the CFO who directs a system that runs forty and flags the three that change the decision. Same title. Completely different leverage on the capital budget.

The Concho example illustrates how consequential development choices can be. Public outcomes alone do not establish what analysis the team lacked. The organizational question is more useful: does the owner of the capital decision also control the quality, timing, and review of the analysis that supports it?

The practical version: energy companies that succeed with AI need someone in the room where capital gets allocated who also controls the AI deployment that feeds that decision. Call it what you want — Chief Decision Officer, VP of Decision Intelligence, the CFO who actually uses the tools. The title doesn't matter. What matters is that the person directing AI and the person signing the AFE are the same person, or sit next to each other. When they're three org layers apart, AI gets aimed at the G&A line. When they're the same person, it gets aimed at the capital budget.

A large capital-to-overhead ratio makes decision support worth examining. It does not establish a return on AI. Give a named decision owner responsibility for the pilot, an explicit quality bar, and a budget for independent review.

DRI case · Treasury and lender readiness

A mid-market E&P with a $1.2B reserve-based lending facility runs two borrowing-base redeterminations per year. Each cycle: 3 analysts, 3 weeks, rebuilding the same variance tables and lender Q&A packets from scratch. The lender asks the same fifteen questions every cycle. Most teams haven't structured last cycle's work product for reuse.

Illustrative pilot design. Start with the recurring packet: source comparison, variance tables, covenant extraction, and lender Q&A. Measure preparation and reviewer hours on completed cases, then evaluate held-out cases with the same acceptance criteria. Price model usage, software, integration, and setup. Any benefit to facility utilization or capital allocation needs separate evidence; it is not included in the labor business case.

What stays human: negotiation posture, lender relationship, representation of downside scenarios, signoff. What the system composes: source comparison, variance tables, covenant extraction, Q&A draft generation, package assembly — delivered before the cycle starts, not rebuilt from scratch. Full beachhead breakdown →

The job that's easiest to automate is almost never the job where automation matters most.

Score by technical exposure and you get one ranking. Score by dollars and you get a different one. That's the whole problem.

The deploy-first paradox
0
roles in the high-value, low-risk corner. In energy, value and consequence always travel together.
TL;DR

There's no free lunch. Every role where AI creates massive value also carries real downside. That makes this a governance problem, not a tech problem.

Exhibit 5 · Value-risk frontier · 30 roles

In energy, the biggest AI upside lives next to real downside.

Everybody wants the lower-right corner of this chart: high value, low risk. Look at it. It's empty. Every role where AI creates massive value also carries real downside. That's not a bug — that's energy. Value and consequence have always traveled together here. The deploy-first corridor isn't where risk is zero. It's where the math works anyway.

X = value creation potential. Y = asymmetric downside risk. Size = automation exposure. The empty lower-right is the chart's most important feature — you can't deploy AI where it's most valuable without accepting real downside. That makes this a governance problem, not a technology problem.

Three dimensions of AI impact · 30 roles

One score hides the answer. Three scores reveal it.

A reservoir engineer barely moves on automation but lights up on value creation. A roughneck scores low on automation but the downside risk is catastrophic. Score on one axis and you automate the wrong things. Score on three and you see which roles are actually worth touching first. These are the 30 we modeled in depth.

From thesis to evidence / A 30-day pilot

Pick one packet.
Make the result count.

Start with borrowing-base preparation, a title exception register, or regulatory exhibits. Choose work that repeats, has a known reviewer, and can be checked against the source.

DAYS 01–05

Freeze the baseline

Choose 10 completed cases. Record preparation hours, reviewer hours, corrections, and cycle time. Agree on material-error definitions before running the test.

DAYS 06–15

Build with five

Use five cases to develop the workflow. Require source references, reconciled totals, and an exception log. Track software and model spend.

DAYS 16–25

Test the other five

Keep the remaining cases unseen during development. Have a domain reviewer compare outputs against the baseline. Count every correction and every minute of review.

DAYS 26–30

Decide with evidence

Expand only if accepted output takes less total human time and costs justify the change. Investigate any escaped material error. Five test cases support a pilot decision, not a reliability claim.

The scorecard: preparation hours + reviewer hours; source-reference accuracy; reconciled material numbers; unsupported claims; escaped errors; software and model costs. Track external approval time separately. A faster draft is useful only if the accepted result is cheaper or better.
Deployment

Start where the result can be measured.

This is not a vision deck. Three workflows you can start next quarter. Inputs named. Gates identified. Compression you can measure. You own the document assembly. The external parties (lenders, counterparties, regulators) own the outcome — but you don't need their permission to start the analytical pilot. Just someone willing to try it.

The gap between "models can do this" and "your company actually does this" is where all the early-mover advantage is. It's closing. These three are where to start — not because they're exciting, but because they can pay for themselves in a single cycle.

Beachhead 01 · Semi-controlled

Start with the packet every lender asks for

Borrowing-base support, covenant monitoring, lender Q&A, amendment packages
What feeds inReserve reports, production exports, price decks, hedge schedules, covenant definitions, prior lender memos
What compressesSource comparison, variance tables, covenant extraction, Q&A draft generation, package assembly
What stays humanNegotiation posture, lender relationship management, representation of downside scenarios, signoff
Success metricsPackage cycle time, analyst hours per cycle, questions answered inside 24 hours, error escapes
Decision contextBorrowing-base readiness · $500M–2B credit facility
The first-principles question: The borrowing-base package exists because lenders can't continuously monitor collateral. If AI makes continuous monitoring cheap, the semi-annual cycle disappears — the lender gets a live dashboard and the "package" becomes unnecessary. You can make the cycle 5x faster. Or you can ask why the cycle exists at all.
Back-of-envelope economics (not audited): 3 analysts × 3 weeks × 2 cycles/year ≈ $35K in allocated labor. AI inference cost ≈ $2K–20K/year. Direct savings: modest. The real value isn't the labor delta — it's the 15 additional scenarios the freed analysts can run, and what that does to the quality of the credit decision on a $500M-2B facility. A 1% improvement in borrowing-base utilization on $1B = $10M.
Beachhead 02 · Semi-controlled

Attack the paper trail that keeps cash and decisions stuck

JOAs, AFEs, JIB review, division orders, curative, title chains
What feeds inJOAs and amendments, lease and deed chains, title opinions, curative notes, division orders, AFEs, election notices
What compressesClause extraction, document comparison, exception queue triage, owner-response drafts, curative clustering
What stays humanLegal interpretation on edge cases, negotiation with counterparties, escalation on title risk
Success metricsSuspense resolution time, title queue aging, JIB dispute cycle time, % auto-triaged
Decision contextNet revenue interest accuracy · $50–500M working-interest exposure
What would prove this wrong: If AI can't reliably parse the clause structure of a 40-page JOA — specifically conditional consent provisions and pooling elections — the comparison step doesn't compress. Current models handle this for standardized JOAs but struggle with pre-2010 custom agreements. Test it on your worst JOA before committing.
Beachhead 03 · Semi-controlled

Shorten the prep stack before the next rate case

IRPs, rate-case exhibits, testimony support, discovery responses
What feeds inForecast workbooks, generation scenarios, depreciation schedules, historical filings, commission discovery requests
What compressesExhibit assembly, testimony draft support with citations, discovery routing, consistency checks
What stays humanPolicy judgment, regulatory strategy, witness preparation, external advocacy
Success metricsTurnaround on data requests, revision count, witness-prep load, consistency escapes
Decision contextAllowed return on equity · $1–10B+ regulated rate base
What would prove this wrong: If a PUC rejects AI-assisted testimony or discovery responses — not because they're wrong, but because the commission won't accept the process. Regulatory conservatism is the binding constraint. Test with a non-controversial filing first, not the general rate case.

Six things AI won't fix. No matter what the vendor tells you.

External clocksAI cannot speed up the regulator's calendar

You can assemble the rate-case exhibits in two days instead of two months. The PUC still takes fourteen months to issue an order. AI compresses time to prepare. Time to decide (agency review, public comment, commission deliberation) runs on a calendar you don't control.

Organizational frictionIt will not stop people from waiting on each other

The borrowing-base takes three weeks because the VP is traveling, the bank wants a different format, and the geologist is arguing with the reservoir engineer about type curves. Maybe 30% of cycle time is work product. The rest is waiting. AI compresses the 30%.

Data qualityIt will not rescue bad source data

Well logs from the 1970s were hand-transcribed onto paper by a guy named Earl. (Seriously. His name was usually Earl.) AI on bad data produces bad results faster. Garbage in, garbage out — now at the speed of light.

IndependenceThe machine is not the signer

Reserve auditors exist because lenders require independent attestation. The bank doesn't care how smart your model is. They care that a human with a PE license signed the page.

Never aloneSome calls stay human because the downside is physical

Final legal opinions, auditor conclusions, safety-critical field calls. These stay human. Not because the industry is slow to change. Because a wellhead blowout doesn't have an undo button.

The downcycleIn a downturn, the value proposition flips

This analysis is implicitly mid-cycle. In a downturn, the prep stack gets gutted by layoffs before AI touches it. The value proposition flips: not "compress the work" but "maintain analytical capability after the RIF." The company that cut 40% of finance in 2020 and has AI can still run the analysis. The one that cut 40% without it can't.

What actually changes inside a role

Inside every role, the same split.

Some tasks shrink. Some disappear. Some become more valuable because the prep bottleneck is gone. The table below decomposes five roles into their task layers — then shows how time and value restructure when the assembly work gets cheap.

Role Tasks AI compresses Tasks AI amplifies Tasks that stay human
Reservoir engineer Decline curve fitting, type-curve generation, reserve report assembly, data gathering from production databases, variance commentary drafts Scenario comparison (can now run 40 instead of 4), sensitivity analysis across price/decline/spacing, pattern recognition across analogue wells Subsurface judgment calls, well spacing decisions, reserve certification signoff, risk framing for the board
Treasury analyst Borrowing-base packet assembly, covenant compliance checks, lender Q&A drafts, amendment redlining, data reconciliation across systems Exception detection (catches covenant breaches earlier), cross-cycle comparison (persistent memory across redeterminations), scenario stress testing on covenants Lender negotiation posture, downside framing, final representations, signoff authority, relationship management
Land / title analyst Division order calculation, lease abstraction, JIB exception screening, curative document comparison, ownership chain reconciliation Cross-asset title pattern matching (flag similar defects across properties), historical exception memory (recalls curative outcomes from prior cycles) Curative negotiation, title opinion judgment calls, counterparty relationship management, legal liability decisions
Regulatory analyst (utility) Exhibit assembly, discovery response drafting, data request compilation, testimony support document preparation, precedent citation lookup Cross-docket pattern analysis (identifies commissioner tendencies), IRP scenario modeling (more alternatives evaluated), consistency checks across multi-year filing history Regulatory strategy, testimony delivery, commissioner relationship management, settlement negotiation, policy judgment
Lineman Paperwork: daily job briefing forms, time entry, incident reporting templates, outage documentation Predictive routing (AI optimizes storm response dispatching), outage pattern recognition (learns from historical restoration sequences) Climbing, switching, grounding, live-line work, safety assessment, crew leadership, storm response decisions

The task examples are qualitative illustrations, not time-and-motion observations. This edition removes the earlier before/after percentages because they were not validated against role-level measurements. Actual time saved depends on source quality, workflow design, and the review required for an accepted result.

The role does not vanish. The economics of the role change.

For a reservoir engineer, faster preparation may release time for interpretation and alternative designs. Whether that makes the work more valuable depends on the quality of those alternatives and the decisions they inform. A compressibility score alone does not establish a percentage of hours saved.

For a treasury analyst, the test is whether a reconciled, traceable lender package takes less total human time to produce. Negotiation and downside framing remain part of the role. Track how released capacity is used before assigning it economic value.

For a lineman, document preparation and dispatch support are a smaller part of a job centered on physical execution and safety. That contrast motivates the benchmark’s layer comparison; it does not measure actual time allocation.

Failure mode to monitor

The "stays human" column is only durable if organizations actually invest the freed-up time in judgment and analysis rather than simply reducing headcount. The risk: a company compresses its five-person treasury team to two people, but keeps the same workflow volume — so each remaining person rubber-stamps AI outputs at 3× speed instead of reviewing them at appropriate depth. The amplification column only works if people have the time budget to fill it. If compression translates to headcount cuts rather than role restructuring, the centaur model degrades into an automation model with a human-shaped rubber stamp at the end. This is an organizational design choice, not a technology constraint.

Across these examples, preparation, decision support, and human accountability are distinct activities. A pilot should measure them separately so faster drafting does not hide extra review or weaken the decision.

Workforce · ScenarioTHREE-PART FRAMEWORK

Work doesn't disappear. It migrates.

"Which jobs does AI replace?" is the wrong question. The right one: what part of the work compresses, what human capability gets amplified, and what new work shows up because AI now exists? The labor evidence so far points toward redesign, not extinction. The functions show up before anyone invents a title for them.

Compresses

The assembly layer inside the job

Drafting, reconciliation, packaging, routing, repetitive evidence assembly. Not "the whole job." The prep work underneath the judgment. Where four analysts assembled borrowing-base packets, two review AI-assembled packets and spend the freed hours on scenario analysis.

Analysts, packagers, coordinators, memo-builders, context routers, reconciliation specialists

Amplifies

Judgment seats get more leverage

Judgment, negotiation, relationship management, signoff, testimony, exception handling, field decision-making, political and regulatory sense-making. These people don't disappear. Their leverage increases because the prep work around them gets cheaper. As analysis gets cheap, judgment gets expensive.

Approvers, negotiators, operators, witnesses, relationship owners, field decision-makers

Emerges

A new control layer around the machine

Every workflow AI automates creates a new control problem. Someone has to teach, audit, route, and sign for machine output. The first new jobs are not sci-fi. They are the people who ensure the machine's work is trustworthy enough to act on.

Evidence architects, exception managers, eval/QA leads, workflow owners, provenance leads, apprenticeship stewards

Don't ask whether AI replaces the job

Ask which part of the seat was the job.

Each tile is a task inside one role. Some flow to AI. Some stay with the human. A few recombine into new functions. The seat is being unbundled — not eliminated.

AI doesn't delete the org chart

It redraws the center.

The old pyramid thins in the middle. The decision layer stays. The field layer stays or grows. A new thin control layer appears around the machine.

The hours don't disappear

They move to higher-consequence work.

Hours leave drafting, reconciliation, and packet assembly. They arrive in exception handling, scenario exploration, and AI governance. The hours don't vanish — they migrate upward.

You can automate the training ground out of existence

Then who makes the decisions in 2035?

Junior manual work builds pattern recognition, which builds judgment, which earns signoff authority. AI slices out the junior layer. The firm now has a problem: how does it produce future decision-makers?

New work functions, not job forecasts

We estimate new work, not new jobs. Functions appear before titles do. The demand model:

Scenario expansionWork generated by running 40 scenarios instead of 4. In treasury, this creates an evidence architect who makes each redetermination cycle cheaper than the last. Exception adjudicationThe 5% of cases that create 95% of delays. In title/non-op, this is the exception manager who handles curative failures that AI flags but cannot resolve. Eval / governanceTrust infrastructure. In regulatory, this is the provenance/QA lead who prevents citation escapes and witness-prep chaos before a rate case. Knowledge maintenanceKeeping the institutional memory alive as senior staff retire. The workflow owner who ensures AI-assisted processes encode domain knowledge, not just patterns. Training / apprenticeshipAcross the company, the apprenticeship steward who keeps junior staff learning judgment, not just becoming prompt operators.

These functions deliver business value, not soft benefits. An evidence architect brings compounding memory. An exception manager brings downside protection. An eval lead brings the trust that lets you actually deploy. A workflow owner brings adoption and budgeted ROI. An apprenticeship steward brings future judgment supply.

The net employment math

A mid-size operator creates 3–7 FTEs across emerging functions while restructuring 20–50 above-field positions. Net headcount likely declines. Per-person value rises. Total payroll may stay roughly flat. The IEA's 1.7-to-1 retirement-to-entrant ratio means in many cases AI-driven compression doesn't cause layoffs — it prevents the capability gap from widening. A 50-person technical team that loses 8 to retirement and replaces 5 (with AI assistance) maintains roughly the same output. HR calls this "transformation." Everyone else calls it Tuesday.

2025–2026: Pilot phase — emerging functions handled as side responsibilities. 2027–2028: Production phase — workflow architect and eval roles become distinct positions. 2029+: Regulatory pressure formalizes governance and provenance roles. The firms that benefit most treat this as workforce restructuring, not simple automation.

Confidence: scenario

These work functions are projected from benchmark structure and cross-industry labor evidence, not observed in operating companies today. The pattern — technology shifts create control work around the seams of the technology — shows up every time (database administrators before databases; DevOps before cloud; compliance officers before Sarbanes-Oxley). We are confident the functions will exist. We are not confident about timing or headcount.

The winners will not be the firms with "fewer people everywhere." They will be the firms with fewer people rebuilding the same packet, more people governing faster decisions, and a deliberate system for producing future judgment.

Decision leverage

Your first workflow depends on what kind of company you run.

Eight company types. Eight different entry points. The shale E&P starts in treasury. The utility starts in regulatory. The PE fund starts in portfolio monitoring. Start in the wrong place and you burn six months building a demo that gets polite applause and zero budget. Each one is a DRI-shaped problem: a single owner, a recurring cycle, a budget line, and company-controlled prep work. No regulator, lender, or board permission needed to start the analytical pilot. Click any row.

Company typeHighest-torque applicationDecision scopeWhy it's asymmetric
Exhibit 6 · Benefit by company type

The same tool can matter more in a different company.

This exhibit is an analyst hypothesis about where strategic value may concentrate. Its efficiency and strategic-value scores are comparative indices, not measured savings or investment returns. Treat them as a way to choose a pilot and a budget owner; use observed workflow results to make the investment decision.

Gray = efficiency gains (cost savings, time compression). Olive = strategic value (decision quality, competitive moat, new capabilities). The gray bars are roughly equal. The olive bars aren't. Scores are modeled estimates — the ranking is the claim, not the precision.

If you're building an AI platform for energy, cost savings aren't your market. Decision leverage is.

The next three exhibits shift the denominator from headcount to token demand, recurring revenue, and capture sequence.

Exhibit 7 · April workflow scenarios · Top 10 of 20

Measure the work before you buy the model.

The April scenario set estimates annual token volume and the share requiring fresh reasoning for 20 workflows. This ranking multiplies those two assumptions. It helps size a test; it is not a forecast of demand, revenue, or value creation. The complete input table follows.

Bars show assumed annual fresh-reasoning tokens per company (total annual tokens × fresh share). “Fresh” is a workload assumption, not a measured cache-miss rate or an automatic model recommendation. No return per token is assigned.

Exhibit 8 · Top 20 recurring workflows

Twenty workflows where AI demand keeps coming back.

Each workflow has a budget owner, recurrence, and an illustrative annual token budget. The April assumptions are kept visible so readers can challenge them. Volume is annual per company and already includes recurrence; do not multiply it by cases per year again. Choose models using task-level quality and cost evaluations, not a job title.

WorkflowCompany typeBudget ownerCases / yrAnnual tokens (M)Fresh %Value at stakeExpansion path
Exhibit 9 · Platform economics · 20 workflows

The model bill is only the first line of the business case.

Price the tokens, then price the work needed to make the output usable. Change the blended rate below to see API costs for the April volume scenarios. Review time, software, integration, and implementation belong in the same business case. Build yours with the pilot calculator →

Annual API cost = annual million tokens × dollars per million. Illustrative blended rates are sensitivities, not vendor quotes. Earlier value-to-token multiples have been removed because their units were inconsistent and the value assumptions were unvalidated.

The loop

Energy is the bottleneck on AI's next wave.

AI compute demand needs chips. Chips need data centers. Data centers need power. Power needs permitting: scenario analysis, regulatory filings, engineering studies, contract negotiation. Twenty workflows stand between a signed lease and a spinning turbine.

The energy connection is real; the causal shortcut is not. At end-2025, the U.S. queue contained 1,312 GW of generation and 749 GW of storage requests. In a separate 2000–2020 entry cohort, 13% of requested capacity had reached operation by end-2025. These are different populations. Neither statistic estimates how much data-center capacity AI-assisted paperwork can unlock. Berkeley Lab, 2026.

The economic loop is visible but undeveloped: faster permitting expands compute supply, which accelerates AI deployment.

The investable question is project-specific: which step is on the critical path, who controls it, and what happens if it finishes earlier? If the turbine delivery date is binding, a faster filing may create no schedule value. If a correct study is the gating item, preparation time can matter a great deal.

A useful system makes the next decision easier to defend: one set of source documents, reconciled assumptions, explicit exceptions, and a named reviewer. Measure the hours released and the mistakes caught before claiming a shorter project schedule.

The most advanced technology on earth is waiting on a filing cabinet.

Energy companies sit on some of the most honest signal in any industry. Every transaction, every filing, every interconnection study, every redet cycle — these aren't survey responses or ad clicks. They're facts about how capital moves through the physical world. The quality of any intelligence layer is only as good as the signal feeding it. Energy has the signal. It just doesn't have the model yet.

AI investment and energy investment increasingly depend on one another. Better preparation can support that relationship. The amount of additional power delivered depends on what actually constrains each project.

What could break this

Faster preparation does not guarantee earlier approval. The previous estimate that 60–70% of queue time was document work has been retired because it was not validated against project-level records. Engineering, equipment, construction, financing, and external review all need to be mapped before assigning schedule savings.

See it happen

Watch AI work a rate-case request.

Illustrative mockup — not a benchmarked model run. A fictional PUC-style data request for illustration. Walkthrough shows plausible sequence, outputs, and speed. No prompt, model, corpus, or evaluation trace is published. Treat it as a storyboard, not evidence.

Docket No. 2024-00187-EL
Staff Data Request Set 3, Item 14
PUC-TX

Request: Provide the Company's actual and projected plant additions, retirements, and transfers for each functional category for the test year and each of the five preceding calendar years. Include explanations for any year-over-year variance exceeding 10%.

Supplemental: For each variance explanation, identify the specific capital project(s), their FERC account classification(s), the date placed in service, and whether the investment was included in the Company's most recent depreciation study filed in Docket No. 2021-00042-EL.

Format: Provide in Excel format with supporting workpapers. Cross-reference to the Company's response to Staff DR Set 1, Item 7 (rate base roll-forward) and OPC DR Set 2, Item 22 (depreciation schedules).

Response deadline: 10 business days from date of service. Objections due within 5 business days.

AI Analysis
Processing document...
Illustrative structured output
What the analyst would do next
What changes the economics? Preparation time saved, review added, and the value of the capacity you can actually use. Test your assumptions in the pilot calculator →
Where you stand

Most companies are earlier than the board thinks.

Use this maturity ladder to describe your own deployment. The dots are illustrative, not a survey of energy companies, and the stages are not a predicted timeline. Progress means producing accepted work with measured quality and cost; buying access to a stronger model does not establish maturity.

Governance

The question is no longer whether to use AI. It is where to trust it.

Match oversight to the consequence of an error and the evidence available about the workflow. Define permission boundaries, deterministic checks, exception handling, and a named accountable reviewer. A model tier is not a substitute for validation; financial transactions and safety-critical decisions require the controls appropriate to the underlying activity.

Tier 1
Bounded automation
PayrollTrade confirmsAP / ARScheduling
Example models (September 2026): Haiku 4.5 / GPT-5.6 Luna · High volume; validated rules and exception handling
Tier 2
Draft & review
Lender packagesCIMsBoard packsContract summaries
Example models (September 2026): Sonnet 5 / GPT-5.6 Terra · Human reviews every output
Tier 3
Assist & decide
Reserve analysisScenario modelingPortfolio optimizationIRP planning
Example models (September 2026): Fable 5.1 / GPT-5.6 Sol · Frontier reasoning, low volume
Tier 4
Human only
Safety-critical opsLegal opinionsAuditor attestationRegulatory testimony
No model. Human judgment, human liability, human signature.

Model tier mapping is a starting point for evaluation, not a performance guarantee. Published API rates below distinguish input from output; a blended rate depends on actual usage. Fable 5.1 cache reads are $0.25 per million tokens. Sol’s $4/$20 rate is promotional, available at least through November 21, 2026. Neither a cache discount nor an API price measures total workflow cost.

USD per million tokens · checked September 7, 2026
ModelInputOutput
Claude Fable 5.1$10$50
Claude Opus 5$5$25
Claude Sonnet 5$2$10
Claude Haiku 4.5$1$5
GPT-5.6 Sol$4$20

Sources: Anthropic pricing and OpenAI Sol model documentation. The $8/M used in cost illustrations is a sensitivity assumption, not a quote.

Everything above: what to deploy now
Everything below: where the industry goes
Implementation

Most energy AI projects die in the same place.

Gartner reported that at least 30% of generative AI projects would be abandoned after proof of concept by end-2025, later revising to at least 50% overall — and that's cross-industry, not energy-specific. In energy, the pattern is consistent: the pilot works, the team gets excited, then data integration hits and the whole thing stalls. The exec who championed it gets a new role. The team quietly shelves it and no one brings it up at the next offsite.

If you've been through a failed "digital transformation" — and most energy executives have — here's why this is different. The last wave tried to change the workflow. New platform, new data architecture, new operating model. Eighteen months of IT integration before anyone saw a result. This doesn't change the workflow. The borrowing-base cycle still exists. The treasury team still runs it. The lender still asks the same fifteen questions. The only thing that changes is that the assembly takes two days instead of three weeks. Nobody adopts a new system. They review AI output instead of building from scratch. Correction, not creation.

That's the structural difference: digital transformation was a platform play — replace the system, retrain the team, migrate the data, hope it works. This is a workflow play — same system, same team, same data, different starting point. Instead of a blank screen, you start with a draft. The analyst who spent three weeks building the packet now spends two days reviewing it. The skills don't change. The starting point does.

The companies that survive this still have to do three things the others don't.

Survival tactic 01

Start with the workflow the team already hates. Borrowing-base assembly, not reservoir analysis. Few will fight to keep a hated process. The treasury team that dreads the three-week redet cycle is your first adopter — they'll champion anything that ends it.

Survival tactic 02

Let the AI draft. Let the human edit and take credit. Measure time saved, not people replaced. The metric is "hours back" not "heads out." The analyst who used to spend three weeks building the borrowing-base packet now spends two days reviewing the AI's draft and a week running scenarios the team never had time for.

Survival tactic 03

Instead of waiting for the reservoir engineer's spreadsheet, send them an AI-drafted variance table built from production data. They'll spend twenty minutes correcting it instead of three days building it. Correction is faster than creation. Use that everywhere.

Scenario · The number not widely reported
$1.0T
in capital reallocation over a decade
Range: $340B (1% improvement) to $1.70T (5%) — we use 3% as the base case

Global energy investment is estimated at $3.4 trillion in 2026. A purely illustrative 3% improvement applied to that base equals $102 billion per year, or $1.02 trillion over ten years at a constant annual investment level. No discounting or adoption curve is applied. That arithmetic shows the scale of the question; it does not show that AI can deliver the assumed improvement.

The native-AI energy company · Speculation

Now imagine the company built after reasoning got cheap.

Snowflake didn't put Oracle in the cloud. Uber didn't put a taxi dispatcher on a phone. The native-AI energy company doesn't make the filing cabinet faster. It doesn't have a filing cabinet. It wouldn't need one.

01 · The decision clock

Development plans become continuous, not annual.

The dev plan is annual because it takes six months to build. The borrowing-base is semi-annual because each cycle takes three weeks. When reasoning is cheap, these become continuous. New KPIs follow: not "how many wells did we drill?" but "how many scenarios did we run before choosing?"

02 · Information asymmetry flattens

The buyer assembles its own reserve estimate before entering the data room.

The seller traditionally knew the asset better than the buyer. AI flattens that. A buyer can now assemble its own reserve estimate from public data. Your information moat gets thinner. Same logic applies to lenders, regulators, and counterparties.

03 · Speed risk

Fast analysis without quality architecture is dangerous.

When the analysis cycle was three weeks, errors got caught in review. At three hours, one bad assumption propagates through 40 scenarios before anyone looks. The companies that skip quality architecture will move fast and break expensive things. Governance is the constraint, not speed.

What remains human-controlled: Field safety. Regulatory judgment. Relationship capital. Physical execution. Liability. The native-AI company removes humans from the assembly line, not from the decisions. Each carries external accountability or physical consequence that no model absorbs. The deeper question: the prep stack exists because humans need information in narrative form. A decision system that operates directly on the data needs no translation layer — no memo, no package, no slide deck. The bottleneck was never the decision. It was the preparation. That distinction reshapes where the ROI model points.

The question underneath everything

What does your company understand that is genuinely hard to understand — and is that understanding getting deeper every day?

If the answer is nothing, AI is cost optimization. Cut headcount, improve margins for a few quarters, get absorbed. If the answer is deep — interconnection queue dynamics, reservoir behavior across basins, regulatory filing patterns, lender decision logic — then AI doesn't just augment your company. It reveals what your company actually is.

This benchmark is either a one-time report or the seed of a world model for energy decision-making. The 404 positions are capabilities. The compressibility scores are a first-pass intelligence layer. The deal flow, the filing history, the cycle data — that's honest signal waiting for a model. The question is whether anyone builds it.

Exhibit 10 · The power bottleneck

A queue is a set of proposals. It is not a delivery schedule.

The current queue and the historical completion rate answer different questions. The stock below is generation plus storage seeking transmission interconnection at end-2025; it excludes data-center load requests. The outcome rate follows an older entry cohort. An interconnection agreement also does not prove that a project is under construction.

Active requests · End-2025 snapshot

2,061 GW

1,312 GW generation + 749 GW storage
About 8,200 projects

549 GW has a draft or executed interconnection agreement. This is not a construction count.

Historical cohort · Entered 2000–2020

13% reached operation

Share of requested capacity operating by end-2025. A historical outcome, not a forecast for the current queue.

Median request-to-operation time exceeded 5 years for projects completed in 2025, in regions with data.

Source: Berkeley Lab, Queued Up 2026. Current stock: end-2025. Historical outcomes: capacity entering in 2000–2020, observed through end-2025. Snapshot totals and cohort outcomes must not be multiplied together.

A delay can defer the cash flows of a power project or computing facility, but nameplate gigawatts alone do not identify that value. Utilization, contract terms, capex, margins, and the actual binding constraint determine the economics. Use project cash flows, not a universal revenue-per-GW shortcut.

331 GW
Disclosed US data-center pipeline

Wood Mackenzie’s Q1 2026 pipeline estimate includes proposed projects. About 40% was in active development. A disclosed megawatt is not an operating megawatt.

Wood Mackenzie · Q1 2026 ↗
2 sides
Load requests and supply requests

A data center requests electricity. A generator or storage project requests interconnection to the transmission system. The Berkeley Lab snapshot covers the latter. The two pipelines cannot be added together or treated as the same queue.

Berkeley Lab · Scope ↗
1 test
Find the constraint that sets the date

Map equipment delivery, engineering, permits, financing, and construction against the commissioning schedule. Measure whether a preparation task actually changes that schedule before assigning a cash-flow benefit to faster drafting.

Use the measurement plan ↓
What the next version models

What hyperscaler power teams actually need to know.

State-level breakdown

In 2023, data centers consumed about 26% of Virginia's electricity supply. Texas gets tens of GW in monthly requests. The queue, the regulatory stack, and the power mix are completely different in each market.

Efficiency crossover

Blackwell is ~4x more efficient per token than Hopper. If chip efficiency improves 4x every 2 years but demand grows 2x, when does efficiency stop outrunning demand? That crossover point is the most important number in the industry.

Bottleneck migration

2023: CoWoS packaging. 2024-25: data center power. 2026+: semiconductor fabs. Bottlenecks move. The analysis needs to show the sequence, not just the current constraint.

Turbine supply chain

If every hyperscaler goes onsite gas, who makes the turbines? GE Vernova, Siemens, Doosan, Wärtsilä, Bloom, Caterpillar. The turbine order book is the leading indicator of data center capacity.

If superintelligence is around the corner

Energy isn't a sector AI impacts. It's the sector that determines how much AI the world gets.

This report covers wave one — the filing cabinet, the prep stack, the back office. But if AGI is close, the binding constraint shifts fast. Away from analyst-hours and toward megawatts, interconnection time, and political permission. The question flips from "which function compresses?" to "which region delivers power fastest?" Here are the four waves, from now to the horizon.

Wave 1
Workflow compression

Filings, packages, contracts, reports. What this page scores. Deployable now.

Wave 2
Continuous capital allocation

Dev plans refresh weekly. Every asset ranked against every alternative. Capital-allocation benefits require evidence beyond a faster analytical cycle.

Wave 3
Energy-system control

AI as the operating system for the grid. The IEA (2025) estimates AI could unlock up to 175 GW of additional transmission capacity from existing lines and save up to $110B/yr in the electricity sector if widely adopted.

Wave 4
Physics

Materials science, fusion, novel reactor design, storage chemistry. What changes when intelligence is no longer the bottleneck — and the physical world still is.

Global data center electricity demand

How much power will AI actually need?

A data center goes up in two to three years. The power plant it needs takes longer. The transmission line takes longer still. The permit to build the transmission line takes longest of all. Energy becomes a speed problem before it becomes a cost problem.

Source: IEA, Key Questions on Energy and AI, April 2026. Historical estimates: 415 TWh in 2024 and 485 TWh in 2025. Projection: about 950 TWh in 2030. Dashed line connects published anchors; intermediate years are not asserted as observations. All data centers are included, not only AI facilities.

Compresses

Analysts, packagers, coordinators, memo-builders. The assembly layer inside every role.

Amplifies

Approvers, negotiators, operators, field decision-makers. As analysis gets cheap, judgment gets expensive.

Emerges

35 new roles. Builders, bridgers, orchestrators. See Exhibit 11 ↓

You can automate the training ground out of existence. If entry-level analytical work disappears, where do future operators, traders, engineers, and executives learn judgment?

In nuclear roles, 1.7 workers nearing retirement for every young worker entering. For grid roles: 1.4× (IEA 2025).

Exhibit 11 · ModeledNEW — 35 emerging roles

35 roles at the intersection of energy and AI.

The 35 roles below are an April 2026 hypothesis about work created where energy and AI meet. They are useful prompts for workforce planning, not a validated count of vacancies or a current salary survey. Leverage, scarcity, and headcount are analyst estimates. Validate demand with actual hiring and deployment data before turning a role into a staffing plan.

Three origin categories. Two scoring axes. The roles that matter sit where leverage is highest and talent is scarcest.

X = leverage (decision value flowing through the role). Y = scarcity (time-to-fill, salary premium, talent pool size). Bubble size = estimated industry-wide headcount demand. Color: ■ Builders (AI-for-Energy) · ■ Bridgers (Energy-for-AI & AI-lab GTM) · ■ Orchestrators (hybrid mutations of existing roles). Upper-right quadrant = the talent war zone: high leverage, high scarcity, where every PE firm, utility, and hyperscaler is competing for the same 200 people. All scores are modeled estimates — the clustering pattern is the claim, not individual precision.

Talent war zone · Upper-right quadrant

Seven roles cluster where leverage exceeds 7 and scarcity exceeds 7: Data Center Energy Lead, Energy Strategy Director (AI Lab), Power Procurement Specialist, Interconnection Queue Manager, AI Safety Engineer (Grid), Enterprise Sales — Energy (AI Lab), and AI-Augmented Reservoir Engineer. Microsoft poached GE's former CFO to run energy strategy. Google hired Duke's Tyler Norris for energy market innovation. Amazon staffed 605 energy positions in a single year. They're not hiring AI people. They're hiring energy people who understand AI. There are maybe 200 of them on earth.

What could break this

These roles assume sustained AI investment and continued energy-AI convergence. If frontier model costs collapse faster than expected, some builder roles commoditize. If hyperscaler power strategies shift to long-term utility contracts (removing the custom procurement layer), bridger demand softens. If AI tools become genuinely autonomous, orchestrator roles shrink rather than grow. The scarcity scores also assume current training pipelines — if universities and bootcamps spin up energy-AI programs at scale, premiums compress within 3–5 years. These are 2026 scores, not permanent conditions.

This report is wave one. The binding constraints shift from analyst-hours to megawatts, interconnection time, and political permission. The beachheads work whether the destination is 3x better or 300x different. Here's what would prove us wrong.

What would prove this wrong

We ran the strongest case against our own thesis. Here's what survived.

Most objections to this analysis make one of four moves. Naming them is not a dismissal — it's an invitation to check whether your own objection clears the bar.

Question drift
We claim "these tasks compress." The objection answers "AI won't replace people." Different question. We agree — the roles persist, the prep work inside them shrinks.
Perfect-proxy fallacy
Compressibility scores are directional proxies, not deployment blueprints. Attacking the proxy for not being a perfect predictor misses the purpose — it ranks where to look first, not what to buy.
Autonomy strawman
We never argue for autonomous AI in critical workflows. The thesis is augmented prep: AI drafts, humans verify and decide. Attacking "unsupervised AI" attacks a claim we don't make.
Edge-case universalizing
Finding one workflow where AI fails and generalizing it to the whole stack. Some workflows won't compress. That's already in the data — the "Six things AI won't fix" section maps exactly where.

The single strongest objection that survives all four filters: compressibility as a single axis is too blunt. A high score tells you the task is compressible — it doesn't tell you whether the organization can adopt it, whether the data exists to train on, or whether the regulatory environment allows it.

We agree. That's why 304 of 404 positions carry three scores, not one: compressibility (can the prep work shrink?), criticality (what breaks if AI gets it wrong?), and reasoning demand (does this need frontier-class models or commodity inference?). If you only look at compressibility, you'll automate the wrong things. The three-axis view is how you avoid that.

What we know, what we believe, and what we're guessing

The scores are hypotheses. Scored by one analyst (Raj Mistry). No inter-rater reliability, no blinded scoring, no adjudication log. Employment counts are formulaic (layer base × tier multiplier), not survey-sourced. Compensation uses six layer-level proxies, not role-specific wages. The directional layer pattern is the claim; individual role numbers are structural estimates. Multi-rater validation is planned for a future edition. CONFIDENCE: medium. Layer rankings are high confidence; individual scores are low.

The ~$1T number is a scenario, not a forecast. $3.4T estimated annual energy investment × 3% decision-quality improvement × 10 years. At 1%: ~$340B. At 5%: ~$1.70T. The 3% is an illustrative sensitivity; this report provides no empirical mapping from scenario count to capital returns. The entire gap between "nice efficiency play" and "industry transformation" rests on whether cheaper analysis leads to more analysis (Jevons) or just faster versions of the same 4 scenarios. We believe Jevons is directionally true. The magnitude is speculative. CONFIDENCE: low. The range ($340B–$1.70T) is the honest answer.

Correction, September 2026: the previous 60–70% document-work share of queue time is retired. Procedure descriptions do not measure elapsed time, and stages can overlap. This report does not have the project-level evidence needed to estimate a causal reduction in commissioning time.

The native-AI energy company is a hypothesis, not a finding. We've identified no venture-backed company that has demonstrated a fully AI-native operating model in energy workflows as of early 2026. The analogies (Snowflake, Uber) are retrospective — we know they worked. Whether the pattern transfers to energy is an open question. CONFIDENCE: speculative.

Known gaps. No competitive landscape mapping (Microsoft Copilot, Palantir, C3.ai, and dozens of startups target overlapping workflows). US-centric regulatory framing (FERC, PUC, NRC, RRC) — international frameworks deserve equal depth. Uneven coverage depth: upstream E&P and utilities are granular, refining and petrochemicals less so, the trading floor alone could be 30+ workflows. And the IOCs are already spending hundreds of millions on AI — this research is most relevant to the companies that haven't started. CONFIDENCE: N/A. Scope disclosures.

Speculation · No precedent exists
A venture-backed, AI-native energy operator is plausible — built not to sell tools to energy companies, but to be one. Capital allocation powered by continuous scenario analysis instead of quarterly human assembly. No venture-backed company has demonstrated this yet. But cheap reasoning, expensive expertise, and recurring decisions are three ingredients that have built category-defining companies in every other industry they've appeared in.
This report is an illustration of its own thesis.

If you're reading this and thinking "my company should be producing this kind of analysis internally" — that thought is the thesis. You're living in it.

~$800
Model cost to produce this page ⓘ
~$450K
Equivalent research team cost ⓘ
1
Analyst with domain expertise ⓘ
90–560×
Model cost leverage ⓘ

373 roles, 24 workflows, 7 artifacts. 11 interactive exhibits. Full dataset, claims ledger, and scoring methodology published. Built with Claude Opus 4.6 and GPT-5.4 Pro — millions of tokens of research, drafting, and iteration. Directed, scored, and verified by Raj Mistry. The models assembled; Raj made the calls. No comparable unaided production process was measured.

Methodology note: the $800 figure reflects direct API billing for the final production sessions only (Claude Opus 4.6 and GPT-5.4 Pro); earlier research, iteration, and discarded drafts are not included. Total model spend including iteration is estimated at $800–$5K, giving the 90–560× range shown above. Labor comparator is a benchmark estimate for a comparable industry research team at market rates — not a quote. Neither figure is audited. Token logs and model-mix breakdown will be published in a future edition.

This report was developed with AI assistance and human editorial judgment. We did not measure a comparable team working without AI, so the production process is not evidence of a particular productivity multiple. The next useful test is a real workflow measured before and after deployment.

Go deeper

Discuss this research with AI

Opens a new conversation with a prompt that links back to this report. The model can read the page and discuss the findings, methodology, or implications for your company.

Chat with Claude Chat with ChatGPT
Sunya Scoop — the 2× weekly energy infrastructure newsletter.
Deals, financings, and AI adoption across energy — for investors, bankers, and operators. 2× per week. No spam, unsubscribe anytime. Privacy policy.
Interactive tool

How AI-exposed is your team?

Enter headcount by layer. The tool applies the April benchmark’s mean layer scores and compensation proxies. It is a screening illustration, not a forecast of time savings or jobs lost.

Your team composition
Field operations
Avg score: 3.5 / 10
Technical / engineering
Avg score: 7.6 / 10
Corporate / admin / finance
Avg score: 8.2 / 10
Advisory / legal / land
Avg score: 7.9 / 10
Capital markets / investing
Avg score: 8.1 / 10
Governance / oversight
Avg score: 7.8 / 10
Your team's weighted AI exposure
0.0
out of 10
0
Total headcount
$0
Est. exposed wage bill
0%
Above-field roles
Pilot economics / September 2026

What has to be true for this to pay?

Enter the work your team actually does. The model values usable capacity released after review and subtracts API, software, and implementation costs. Every input is yours to change.

Illustrative first-year capacity value after costs
$1,425

240 hNet human hours released
$960Annual API cost
$8,960Total first-year costs
$10,385Usable capacity value
124.2 hProductively used hours to break even
1.2×Capacity value / first-year costs

Scenario, not a forecast. Released time has economic value only when it is productively used or spending is avoided. No gain from better capital allocation, faster permitting, or fewer errors is assumed here. Those benefits require separate evidence.

Evidence / benchmark vintage: April 2026

Evidence room — full dataset, claims ledger, and scoring method.

The September source ledger and corrections accompany the unchanged April role dataset. Use the reproduction script to check the arithmetic, and the pilot scorecard to collect the evidence this benchmark does not yet contain.

A note on the data. 60 of 404 positions use BLS/industry employment anchors, including analogs rather than exact occupational matches. The remaining 344 use a formulaic estimate: layer base × tier multiplier (three bands: 0.75×, 1.0×, 1.25×). These are structural scaffolding for directional wage-pool analysis, not census-quality headcounts. Compensation uses six layer-level proxies ($58K–$180K), not role-specific wages. The directional finding — that roughly nine of every ten AI-exposed wage dollars sit above the field — holds across a wide range of employment assumptions. Individual role numbers should not be cited as precise. Full methodology: methodology FAQ.
September edition materials

The role dataset is the dated April benchmark. This edition adds external evidence and improves tools; it does not claim newly observed workforce productivity.

How to reproduce every number in this benchmark

The core benchmark — scores, employment, compensation, and the ~90% headline — can be reconstructed from the published CSV and the three formulas below. Some interactive exhibits (role explorer, workflow economics, demand scenarios) use additional embedded data structures visible in the page source.

Formula 1 — Exposure-weighted wage bill

EWB = estimated_employment × layer_median_comp_proxy × compressibility_score / 10

This exposure metric drives the wage-bill ranking, slope chart, and ~90% headline. Heatmap and treemap area use employment, not exposure-weighted wages. The core exhibits trace back to this formula applied row-by-row across the 404-position CSV.

Formula 2 — Employment estimation (344 formulaic rows)

employment = layer_base × tier_multiplier

Where:

layer_basePhysical: 22,000 · Technical: 8,000 · Corporate: 10,000 · Advisory: 7,000 · Capital mkts: 5,000 · Governance: 3,000 tier_multiplierThree explicit bands: Low (0.75×), Mid (1.0×), High (1.25×). Each role is assigned to a tier by a stable hash of its name: h=0; for each char c: h=((h<<5)-h+charCode(c))|0; tier=|h|%3. Tier 0→0.75×, 1→1.0×, 2→1.25×. This produces three discrete employment levels per layer rather than pseudo-random integers that would imply false precision. The tier assignment is deterministic and reproducible from role name alone.

Methodology note (v32.1): Employment estimation is decoupled from compressibility scoring. An earlier version used max(0.15, 1.3 − score/8) as a score factor, which gave low-scoring roles more people by construction. The continuous jitter 0.7 + (hash mod 600)/1000 was also replaced with three explicit tier bands to eliminate false precision. All employment figures for formulaic rows are displayed with rounding and a "~" prefix to signal estimation uncertainty.

The 60 anchor rows override this formula entirely — their employment figures come from BLS Occupational Employment and Wage Statistics or industry-derived analogs, flagged in the employment_source column. Note: two of these 60 rows are not people-roles (Board pack production, A&D diligence support) — they are workflow/artifact rows with industry-estimated volume, not BLS occupational matches.

Formula 3 — Compensation proxies

Physical
Technical
Corporate
Advisory
Capital Mkts
Governance
$58K
$95K
$72K
$105K
$130K
$180K

These are BLS OEWS-derived median annual wages for the layer's representative occupations, not role-specific compensation. Every row in the same layer uses the same proxy. Figures reflect wages only; employer-paid benefits are not included.

Compressibility score — what it measures

Each position's compressibility score (1–10) represents how plausibly current AI can compress the preparation stack around that role or workflow. It is built from five sub-dimensions:

Task exposureWhat fraction of the role's tasks involve document-heavy, pattern-matching, or analytical work that current models handle well Adoption feasibilityHow quickly the workflow could realistically adopt AI given regulatory, organizational, and technical constraints Economic compressionWhether compressing this workflow changes the cost structure enough for a budget owner to notice Workflow standardizationHow repeatable and templated the work is across cycles and counterparties Template densityHow much of the output follows a known structure (forms, packets, exhibits, standard reports)

The score is a calibrated composite, not a simple average. Roles with high task exposure but low adoption feasibility (e.g., safety-critical field roles) score lower than roles where both align (e.g., document assembly and reconciliation). The full scoring rationale is published per-row in the CSV's rationale column and in the scoring method note.

Augmented dataset

404 positions with compressibility scores, employment estimates, and augmented axes. 60 BLS/industry anchor rows, 344 formulaic. The raw material behind every exhibit.

Download 404-position CSV →

Claims ledger

Every major claim tagged by type: measured, modeled, scenario, hypothesis, cited, or self-critique.

Download claims ledger CSV →

Scenario table

Base, aggressive, and extreme annual token scenarios across six workflow families.

Download token-demand scenarios →

Headline metrics

Registry of key metrics with definitions, denominators, and derivation paths.

Download metrics registry →

Scoring method

The five sub-dimensions, calibration approach, and heuristics behind every compressibility score.

Read scoring methodology →

Sensitivity analysis

Monte Carlo stress test: the 90% finding survives ±1 and ±2 point scoring errors across all 404 rows. 1,000 iterations.

Read Monte Carlo stress test →

External sources

Current primary-source register, plus a clearly dated archive of earlier source notes.

Read source validation table →

Counterforces

Where the thesis bends, slows, or fails — and how to test each assumption.

Read counterforces →

The useful next step is small enough to measure: one recurring workflow, one accountable owner, and an agreed definition of an accepted result. The dataset and assumptions are open for inspection.

Start with ten historical cases. Make the sources traceable. Count the review time. Expand only when the evidence earns it.

If you run an energy operator

Name the five decision loops where your team still rations analysis. Start where the cycle recurs quarterly, the team dreads the assembly, and you own the prep — no regulator permission needed.

See the three beachheads

If you invest in or lend to energy

Ask your portfolio companies how many scenarios they ran before their last billion-dollar commitment. Ask which assumptions changed the decision, how alternatives were challenged, and whether anyone tested the analytical process. The gap between 4 and 40 is where the next write-down hides.

Download the dataset

If you build AI tools for energy

The workflow illustrations connect annual token assumptions to recurring work and a budget owner. Start where an accepted output and its full costs can be measured within one operating cycle.

Open claims ledger

Four scenarios on a billion-dollar decision. Or forty. That's the only question this page asks.

Know someone who should see this?

Forward the research to a colleague. The link includes the full dataset, methodology, and proof trace.

Forward to your team →
Download PDF Send feedback

Research should make a decision easier to examine. Challenge the assumptions, test a workflow, and share what you learn. — Raj