A review for the Sutro energy effort · September 2026

Brain vs H100: the physical ledger

Transistors against neurons, switches against spikes, joules and bits against metres. The brain is not cheaper per spike: one spike costs about a hundred thousand transistor switches. It wins by computing below the spike, in protein switches that cost about a thousandth of a transistor's; by firing almost never; by delivering each spike to thousands of weights that never move; and by moving a thousand times fewer bit-metres a million times slower.

80 Bvs86 B
transistors vs neurons; the brain also has 150 T synapses, about 2,000 per H100 transistor
700 Wvs20 W
module TDP vs whole-brain glucose; only 6 W of the 20 W becomes ATP, and 0.1 W of it is cortical computation
4×10¹⁸vs10¹¹
transistor switches per second vs action potentials; synaptic events reach 10¹⁴/s
1.4 pJvs0.13 pJ
per FP16 multiply-add at TDP vs per synaptic event at 20 W: a 10× gap per operation, not a million

Divide each machine's power by what it does and the per-operation gap is ten to a hundred times. The million-fold figure in circulation compares 20 W with a digital simulation of a brain-sized spiking network, which is a statement about von Neumann overhead, not about device physics. Below, every quantity is placed on the same yardstick, with its source and its range, and a section applies the ten principles of Sterling and Laughlin's Principles of Neural Design to both machines.

The ledger

Same yardstick for both machines

Tags: P published, E derived on this page, M measured in sutro-problems, R Rao's claim, H Horowitz's claim. Ranges are the spread across sources; the brain's rows swing with the assumed mean firing rate (0.16–4 Hz), which the literature has not settled.

QuantityNVIDIA H100 SXM5Human brainRatio
ElementsP80 B transistors (7% in SRAM cells); 132 SMs; 2.7×10⁵ FP16 multiply-add units86 B neurons (16 B cortical, 69 B cerebellar); 1.5×10¹⁴ neocortical synapses; 85 B gliabrain: 2,000 synapses per transistor
SizeP814 mm² die, ~0.6 cm³ of silicon; 55×58 mm package; 150×80 mm module1.2–1.4 L; 1.4–1.5 kg; 475 cm³ cortical grey, 451 cm³ whitebrain 2,000× the volume
PowerP700 W module; 1,275 W per GPU in a DGX H10020 W glucose (16 W by O₂, 17 W by PET, 25 W by neuron scaling; Sterling & Laughlin close on 12 W); 2% of body mass, 20% of its energy; 6.2 W as ATP; cortex 4.9 ATP-W; the Na⁺/K⁺ pump takes 60%H100 35× (110× against ATP-watts)
Power densityE1,200 W per cm³ of silicon (86 W per cm² of die)15 mW per cm³, and pumps set the ceiling: one open Na⁺ channel needs 10,000 pump molecules to keep up, cutting sustainable information rate ~10⁵× below channel packing≈10⁵× per unit volume
Event rate per elementP1.98 GHz clock, activity ≈0.1: ~2×10⁸ transitions per gate per second0.16 Hz (energy budget, Lennie) to 1–8 Hz (electrodes) per neuron; power density caps the cortical mean below 10 Hz; 2 Hz costs as much as staying at rest; retinal ganglion cells 4–8 Hz10⁷–10⁹×
Signal swingP0.75 V core supply (0.7–0.9 V range)110 mV action potential; 16 mV from reset to threshold; kT/q = 26 mV floorvoltage² favours the neuron 46×
Capacitance charged per eventP0.1–0.35 fF per switching gate incl. local wire50 pF per 5.5 mm of 0.3 µm axon; ~750 pF per cortical neuron (1 µF/cm²)neuron 2×10⁶× larger
CV² per eventE56.2 aJ–197 aJ per switch9.08 pJ to charge 750 pF through 110 mV; pumping the charge back costs 2–4× thatspike ≈10⁵ switches
Events per secondE4×10¹⁸ switches (2×10¹⁸–10¹⁹); 10¹⁵ FP16 FLOP; 5×10¹⁴ multiply-adds9×10¹⁰ spikes (1.4×10¹⁰ at 0.16 Hz to 3.4×10¹¹ at 4 Hz); 1.5×10¹⁴ synaptic arrivals (2.4×10¹³–6×10¹⁴), a quarter of them releasing a vesicleswitches vs spikes 4×10⁷; switches vs synaptic events 3×10⁴
Energy per event, bottom-upPswitch 56.2 aJ–197 aJ; 8-bit add 30 fJ (45 nm); INT4 multiply-add 21 fJ (5 nm research chip)vesicle release 16.4 fJ (1.64×10⁵ ATP); postsynaptic event 13.7 fJ; action potential proper 38.4 pJ (3.84×10⁸ ATP, rat); human spike with all its synaptic consequences 240 pJ (2.4×10⁹ ATP)per synaptic event ≈ 100 switches
Cheapest switching elementPCMOS gate switch, 56.2 aJ–197 aJ = 1.3×10⁴–4.6×10⁴ kT; switches in picosecondsa protein changing shape: a G protein registers a bit for 25 kT (107 zJ, one ATP), 35× the Landauer floor; an ion channel's gate also moves for ~25 kT; microseconds to milliseconds, and chemistry reaches only ~1 µm per msprotein 526–1,840× cheaper, far slower and shorter-range
Price of going electrical, then pulsedPone mode: full-swing digital at 0.75 V everywhere; analogue only in the I/O linksa Na⁺ channel open 1 ms admits 6,000 ions, and pumping them back costs 2,000 ATP (200 aJ), 2,000× a G-protein cycle; a 1.5 mV graded potential costs 100× less per unit length than a 150 mV spike; spikes carry < 500 bit/s where graded signals carry > 2,000a millisecond of channel current costs about one CMOS switch; so the brain uses chemistry near, graded voltage within ~1 mm, spikes only beyond
Energy per event, all-in (power ÷ events)E175 aJ per switch; 707 fJ per FP16 FLOP; 1.41 pJ per multiply-add; 354 fJ per FP8 FLOP232 pJ per spike (60 pJ–1.4 nJ); 133 fJ per synaptic event (30 fJ–0.8 pJ)multiply-add vs synaptic event: ≈10×
Share of energy spent moving dataPHBM ≈120 W of 700 W (17%); >50% of die energy in caches and register filescommunication 3.5 W vs computation 0.10 W in cortex (35:1); white matter 71% of the cortical budget; by Sterling & Laughlin's accounting spikes are only ~20% of cortical energy, but wire fills ~3/5 of grey matterboth are communication-bound; the brain as much in volume as in joules
Bits moved per secondPregister file 4×10¹⁵; shared memory 2.7×10¹⁴; L2 7×10¹³; HBM 2.7×10¹³; NVLink 7.2×10¹²cortex spike traffic 1.6×10¹⁰ (1 bit, 1 Hz) to 7.7×10¹⁰ (Levy); capacity 10¹²; senses in 10⁹; behaviour out 10HBM vs cortex traffic 1,700×
Energy per bit movedPon-chip 0.1 pJ per bit-mm; HBM 4 pJ; NVLink-C2C 1.3 pJ; long-reach SerDes 3–4.5 pJone fly synapse 5 fJ (5×10⁴ ATP) per bit at 55 bit/s; the 1,320-synapse array it actually runs 200 fJ (2×10⁶ ATP) per bit at 2,100 bit/s; spike coding 0.1–1 pJ; a bit held by a whole neuron ~10⁹ kT (4.28 pJ); whole brain per cortical bit 1.23 nJper bit the brain is not cheaper, and its cost per bit climbs with rate
WireP≈15 metal layers; 1–3 m NVLink copper cables176,000 km of myelinated axon; 1.5–2×10⁶ km of fine axon in all (1.5×10⁶ km for the whole nervous system per Sterling & Laughlin; 4 km per mm³ of cortex gives 1.9×10⁶); finest axons 0.1–0.2 µm; dendrites + axons ≈3/5 of grey matterbrain: 10⁹ m of wire in a litre
Signal velocityP1.25×10⁷ m/s on-chip (80 ps/mm, c/24); 1.4×10⁸ m/s on a board0.5–2 m/s unmyelinated; 3–20 m/s callosal; up to 120 m/ssilicon 10⁶–10⁸× faster
Distance from Landauer (kT ln 2 = 2.9 zJ)Pswitch 2×10⁴–7×10⁴; FP16 multiply-add 5×10⁸; HBM bit 10⁹protein switch 25 kT, 37× the floor; synaptic bit 3×10⁵–2×10⁶; vesicle 6×10⁶; spike 10¹⁰–10¹¹ kT; cortical neuron 10⁸ below the bits-per-joule limitthe brain's chemistry is near the floor; its spikes, and every silicon operation, are not

Swipe sideways to see all four columns.

Switches and spikes

The chip switches 10⁷–10⁸ times more often than the brain fires

Events per second (log scale, per second)

A transistor switch and a spike are not the same amount of work; the right pair to compare is a synaptic event and a multiply-add
H100 transistor switches18.5 B gates × activity 0.1 × 1.98 GHz; power-side bound agrees · range 2×10¹⁸/s–10¹⁹/s
H100 FP16 multiply-addsdense tensor peak, 989.5 TFLOPS
H100 bits out of HBM3.35 TB/s
H100 bits over NVLink900 GB/s
Brain synaptic arrivals1.5×10¹⁴ synapses × 0.16–4 Hz · range 2.4×10¹³/s–6×10¹⁴/s
Brain vesicle releasesa quarter of arrivals release · range 6×10¹²/s–1.5×10¹⁴/s
Brain action potentials86 B neurons × 0.16–4 Hz · range 1.4×10¹⁰/s–3.4×10¹¹/s
Cortex bits per secondRao 16 G to Levy 77 G; capacity 1 T · range 1.6×10¹⁰/s–10¹²/s
Human behaviourZheng & Meister: 10 bit/s
10⁰10³10⁶10⁹10¹²10¹⁵10¹⁸
H100 / siliconbrainreferencelight band = range across sources

Even the brain's most generous count, six hundred trillion synaptic arrivals a second, is under a thousandth of the chip's switching. But a synaptic event is a multiply-add: a weight applied to an incoming spike, with the weight stored in place. On that pairing the brain does 10¹³–10¹⁴ operations a second against the H100's 5×10¹⁴ multiply-adds, at 20 W against 700 W.

Joules per event

A spike costs a hundred thousand switches, a synapse a hundred, and a protein switch a thousandth of one

Energy per elementary event (log scale, joules)

Bottom-up biology at 10⁻¹⁹ J per ATP (36–60 kJ/mol spans ±30%); silicon from published circuit figures; all-in rows divide power by event rate
Landauer limit, 300 KkT ln 2
Protein switch: a G-protein bit, a channel's gate25 kT, one ATP; Sterling & Laughlin
Transistor switch, CV²0.1–0.35 fF at 0.75 V · range 56.2 aJ–197 aJ
Na⁺ channel open 1 ms, ions pumped back6,000 Na⁺, 2,000 ATP; 2,000× the chemical signal
8-bit integer addHorowitz 2014, 45 nm
INT4 multiply-add, research chipNVIDIA 5 nm, 95.6 TOPS/W
Synaptic vesicle release1.64×10⁵ ATP, Attwell & Laughlin · range 13.7 fJ–27.9 fJ
Synaptic event, all-in20 W ÷ arrivals per second · range 33.3 fJ–833 fJ
One bit moved 1 mm on-chipDally: 100 fJ per bit-mm
FP16 multiply-add, H100 all-in700 W ÷ 4.9×10¹⁴ per second
One HBM bitO'Connor 3.97 pJ; measured 4.5 pJ · range 3.9 pJ–4.5 pJ
Action potential, sodium pumping only3.84×10⁸ ATP, rat cortex · range 23 pJ–38.4 pJ
Human cortical spike with its synaptic consequences2.4×10⁹ ATP, Lennie 2003 · range 144 pJ–240 pJ
DRAM access, 64-bitHorowitz 2014 · range 1.3 nJ–2.6 nJ
10⁻²¹10⁻¹⁸10⁻¹⁵10⁻¹²10⁻⁹
H100 / siliconbrainreferencelight band = range across sources

The voltage lever is real but small

E = CV². Against a 0.75 V supply the neuron's 110 mV swing buys 46×; against 1 V, 83×. The capacitance it charges is 2×10⁶× larger, so per event the neuron spends about 10⁵× more than a gate. CMOS cannot follow the neuron down: the energy-optimal supply is 250–350 mV and saves about 10× at 10× the delay; the functional floor is 36–48 mV. Below the spike the brain pays per ion, so cost falls in proportion to swing: a 1.5 mV graded potential costs 100× less per unit length than a 150 mV spike.

The activity lever is enormous

A logic gate transitions about 2×10⁸ times a second; a cortical neuron fires 0.16–4 times. That is 2×10⁸× in favour of the brain, and it is imposed by the energy supply: Lennie's budget sustains 0.16 spikes per neuron per second; 1.8 Hz would spend the whole brain's supply. Fewer than 1 in 6 neurons can be active at once. It is also the optimum: for a retinal ganglion cell to double its bit rate would take 10× the spikes and 100× the space and energy.

Where the efficiency comes from

Six levers, five of them for the brain

LeverFactorFavoursSource of the number
Voltage swing squared46×brain(0.75 V / 0.11 V)²; Horowitz's P = αCV²f
Capacitance per event2×10⁶×chip750 pF neuron vs 0.2 fF gate
Events per element per second2×10⁸×brain1.98 GHz × 0.1 vs 1 Hz
Work per event~7,000×brainone spike is delivered to ~7,000 synapses, each an in-place multiply-add; the chip pays ~3,700 switches per FLOP
Data movement1,700× fewer bitsbrainHBM 2.7×10¹³ bit/s vs cortex 1.6×10¹⁰; but 35:1 communication-to-computation in cortex too
Computing below the spike526–1,840× per switchbrain25 kT protein switch vs CMOS gate (Sterling & Laughlin); not counted in the net row, whose denominator is synaptic events only
Net, per operation10–100×brain0.13 pJ (all-in) or 14 fJ (bottom-up) per synaptic event vs 1.4 pJ per multiply-add

Multiply the levers and the physics gives the brain one to two orders of magnitude per operation. That is a floor for the brain, not a ceiling: the ledger counts synaptic events as its only operations, while the filtering, subtraction and gain control it does in chemistry and graded voltage ride free inside the 20 W. The other four orders in the popular "million times" are the cost of simulating sparse, event-driven, in-memory computation on a clocked machine that fetches every weight: the same cortical microcircuit costs 74 nJ per synaptic event on an RTX 4090 against 19–760 fJ in tissue.

Sterling & Laughlin, Principles of Neural Design (2015)

Information costs more the faster it flows, so the brain keeps every channel slow

Their one-line diagnosis: as information rate rises, costs rise disproportionately. Signal-to-noise improves as √n in the events averaged while cost grows as n; information grows as log(S/N) while cost grows as (S/N)². So the brain builds many slow, matched channels rather than a few fast ones, and the activity lever above is the optimum of a measured cost curve, not a handicap.

Bits per spike fall as the rate rises

Guinea-pig retinal ganglion cells (Koch et al. 2006), in Sterling & Laughlin fig. 11.25
Local-edge cell, 4 Hz140 spikes per 300 bits
Brisk-transient cell, 8 Hz170 spikes per 300 bits
Hypothetical high-rate cell, 40 Hz270 spikes per 300 bits
00.511.522.5
measuredhypothetical

Energy per bit (log scale, joules)

A fly's synapses pay 40× more per bit in the array than alone; a bit held by a neuron costs 4×10⁷ times one held by a protein
One fly synapse, alone55 bit/s; 5×10⁴ ATP per bit
H100: one bit moved 1 mm on-chipDally: 100 fJ per bit-mm
The 1,320-synapse array a fly photoreceptor drives2,100 bit/s: 40× the bits at 50× the cost per bit
H100: one bit out of HBMO'Connor 3.97 pJ; measured 4.5 pJ
A bit held by a whole neuron~10⁹ kT, against 25 kT in one protein
10⁻¹⁵10⁻¹⁴10⁻¹³10⁻¹²10⁻¹¹
H100 / siliconbrainreferencelight band = range across sources

The ten principles, scored against the H100

PrincipleWhat the brain does (Sterling & Laughlin, 2015)What the H100 does
Compute with chemistryA G protein registers a bit for 25 kT, 35× the Landauer floor; a channel's electrical cycle costs 2,000× more ATP. In a rod, the electrical side costs 100× the chemical amplifier (pp. 123, 142, 212)Every operation is electrical: 1.3×10⁴–4.6×10⁴ kT per gate switch
Compute directly with analogue primitivesA binding curve or an RC membrane adds, takes logs, thresholds and multiplies in one step; a digital multiply of b bits takes about b^1.585 steps, 9 at 4 bits (p. 131)FP16 multiply-add built from thousands of transistor switches (this page: ~4,000 per FLOP)
Combine analogue and pulsatileGraded voltage within ~1 mm, spikes only beyond. Recoding to spikes cuts a neuron from > 2,000 to < 500 bit/s and can lose 90% of the information (p. 174); the retina recodes three times before the optic nerveDigital end to end; analogue survives mainly in the SerDes, memory PHYs and clocking
SparsifySpikes cost 1,000× the chemistry that triggers them, so they come rarely, in bursts, in few cells (p. 437); for a cortical neuron 2 Hz costs as much as staying at rest (p. 342)Dense tensor cores, ~10% of gates switching every cycle; 2:4 structured sparsity gives at most 2×
Send only what is neededThe fly lamina subtracts the 75% of its input that is redundant before transmitting it, a 4× saving (pp. 249–252); by the time a ganglion cell fires, 90% of what entered at the cones is gone (p. 393)Re-reads every weight from HBM on every pass: 2.7×10¹³ bit/s, ~17% of module power
Send at the lowest acceptable rateBits per spike: 2.1 at 4 Hz, 1.1 at 40 Hz (p. 320); doubling a ganglion cell's bit rate takes 10× the spikes and 100× the space and energy (p. 393); matching a Drosophila photoreceptor to its 200 bit/s optics makes it 6× more efficient than a blowfly's 1,000 bit/s one (p. 228)1.98 GHz clock and full-rate links, fixed by the design rather than by the data
Minimize wireWire fills ~3/5 of grey matter; branching axons, 3-D dendrites and spines reach the same synapses in 1/45,000 of the volume (pp. 367, 383–384); conduction is 10⁸× slower than copper, so the only economical way to save time is shorter wire (p. 55)Weights travel millimetres at 0.1 pJ per bit-mm; more than half the die's energy goes to caches and register files
Make components irreducibly smallAxons stop at 0.1–0.2 µm: any thinner and one thermally opened Na⁺ channel fires a spike (pp. 175, 185); 30 nm vesicles; ~100 nm³ proteinsAlso at its floor (TSMC 4N). The one lever that favours the chip: 2×10⁶× less capacitance per event
Complicate~20 ganglion-cell types, rods for starlight and cones for daylight; the fly's one-type photoreceptor spends 10× a cone's energy at the same ~200 bit/s (pp. 224–225)132 identical SMs; one tensor-core design for every layer and every input
Adapt, match, learn and forgetCapacity matched to need at every stage; memory sits at the synapse that computes, so recall needs no wire (p. 404); storage costs grow as (S/N)² while capacity grows as log(S/N), so many small noisy synapses win (p. 411); no learning without forgetting (p. 430)Fixed precision and clock whatever the input; weights stored millimetres from the arithmetic

Swipe sideways to see all three columns.

By the book's own account, two of the ten came from electronics (analogue primitives and mixing analogue with pulses) and two are standard chip practice (minimize wire, make components small). The H100 follows the chip-practice pair. The other six, and above all send at the lowest rate and store where you compute, are where an energy-first design still has room.

Horowitz, checked

The energy equation holds; the neuron heuristic is not his

HeuristicRecomputedVerdict
"Computing's energy problem" is ISSCC 2012 HISSCC 2014, plenary 1.1, pp. 10–14, DOI 10.1109/ISSCC.2014.6757323. The 2012 paper is CPU DB (CACM), its reference 5.off by 2 years
45 nm table: 8-bit add 0.03 pJ … DRAM 1.3–2.6 nJ HReproduced from the printed figure. The slides differ on three multiplies (3 vs 3.1 pJ, 1 vs 1.1, 4 vs 3.7); prefer the paper.holds
Neurons at 100 mV, CPUs at 1 V, so 100× less power SThe arithmetic follows from CV²; against the H100's 0.75 V it is 46×. Not in the paper or its 46 slides (no brain, neuron or millivolt-as-neuron anywhere). Silicon cannot realise it: near-threshold saves ~10× at ~10× the delay (Dreslinski 2010).roughly, unsourced to Horowitz
Neurons accept leakage to run that low SResting potentials take 13% of grey-matter signalling energy (Attwell & Laughlin), 28% + 10% glial of total cortical energy (Lennie); the kT/q floor sets the on-off ratio in both substrates.holds
Overhead, not arithmetic, is the cost H70 pJ per instruction around a 0.1 pJ add; 30 pJ fetch/decode on the H100's 1.5 pJ FP16 FMA (2000%), 22% on a matrix instruction. The brain has no fetch or decode.holds
Information

One H100 moves 1,700× the bits of a cortex; which cortex number depends on the cut

Bits per second (log scale)

Ten orders of magnitude separate what the senses gather from what behaviour emits; the GPU numbers are datasheet or microbenchmark
H100 register filederived from the FP32 datapath
H100 shared memory128 B/clk/SM measured
H100 L24,472 B/clk measured
H100 HBM3Rao: “nearly 30 trillion”
H100 NVLink900 GB/s
Cortex, axonal capacityLaughlin & Sejnowski: 100 bit/s per axon
Cortex, spike trafficRao 16 G (1 bit, 1 Hz) to Levy 77 G · range 1.6×10¹⁰/s–7.7×10¹⁰/s
Corpus callosum2×10⁸ axons; traffic to capacity · range 3.2×10⁷/s–2×10¹⁰/s
Senses inZheng & Meister: cones alone 1.6 Gbit/s per eye
Retina outKoch 2006: ~10 Mbit/s per eye
Behaviour out10 bit/s
10⁰10³10⁶10⁹10¹²10¹⁵
H100 / siliconbrainreferencelight band = range across sources
Rao's claim (All-In Summit, Sept 2026)RecomputedVerdict
The cortex moves about 16 billion bits a second R16.3×10⁹ cortical neurons × 1 spike/s × 1 bit = 16.3 Gbit/s. A traffic estimate: at Levy's 5.1 bit/s per neuron it is 77 Gbit/s; the axonal capacity is ~1 Tbit/s. Sensory input is ~1 Gbit/s and behaviour 10 bit/s.roughly, ±1 order
A GPU moves nearly 30 trillion bits a second in and out of memory R3.35 TB/s × 8 = 26.8 Tbit/s. A100-80GB: 16.3 Tbit/s.holds
Inside the chip 10–100× more RL2 2.6×, shared memory 10×, register file ~160× the HBM rate.holds
The brain computes a million times more efficiently (Sequoia AI Ascent, May 2026) RTrue for the same job: 74 nJ per simulated synaptic event on an RTX 4090 vs 19–760 fJ in tissue. Per device event it is 10–100×; per bit moved the brain pays more (1.23 nJ vs 26.1 pJ with all power charged to traffic).mixes levels
Bits over distance

Copper is not dearer than axon per bit-metre; the brain just moves 1,000× fewer of them, a million times slower

LinkBandwidthLengthSignal velocityEnergyJ per bit-metrebit-metres per second
0.3 µm unmyelinated cortical axon8 bit/s (4 Hz, 2 bit/spike)~1 mm0.5–2 m/s0.26 pJ per spike-mm1.3×10⁻¹⁰0.008
1 µm myelinated axon8 bit/s10 cm≈6 m/s0.058 pJ per spike-mm2.9×10⁻¹¹0.8
Corpus callosum, 2×10⁸ axons1.6 Gbit/s (0.03–20)10–15 cm3–20 m/s; 12–24 ms across—2.9×10⁻¹¹1.6×10⁸ (5×10⁶–3×10⁹)
Whole brain, all axons10¹⁰–10¹² bit/s1.76×10⁸ m myelinated + ~10⁹ m cortical0.5–120 m/s0.2–0.7 W in total~10⁻¹⁰~5×10⁹ (10⁹–10¹⁰)
H100 on-die global wire—10 mm1.25×10⁷ m/s (80 ps/mm)1 pJ per bit per 10 mm1.0×10⁻¹⁰L2: 7×10¹¹
H100 HBM32.7×10¹³ bit/s~5 mm interposer (+10 mm inside the stack)~10⁸ m/s4 pJ per bit7.8×10⁻¹⁰1.3×10¹¹
H100 NVLink 47.2×10¹² bit/s1–3 m copper1.4×10⁸ m/s; 7 ns per metre1.3–5 pJ per bit1.3–5×10⁻¹²7.2×10¹²

Three normalisations, three different winners. On latency silicon wins by 10⁶–10⁸. On throughput × distance one H100 carries about a thousand times the bit-metres of a whole brain, and thirty times more per watt. On energy per bit-metre the two substrates are within a factor of four of each other, and a long-reach serial link beats both. Myelin buys speed, not energy: it only saves ATP above 12 Hz, and cortical axons fire at 4. The brain's edge is architectural, not material: a spike travels micrometres to reach thousands of weights, where the chip fetches every weight across millimetres.

Which number to quote

Six brain-vs-computer comparisons and what each actually compares

ComparisonGapWhat is being compared
20 W brain vs 7.9 MW digital simulation of a brain-sized spiking network (Mehonic & Kenyon, Nature 2022)4×10⁵doing the brain's job on a von Neumann machine: memory traffic, float precision, time-stepping quiescent synapses
74 nJ per synaptic event on an RTX 4090 vs 19–760 fJ in tissue (Senk et al. 2025)10⁵–10⁶the same cortical-microcircuit model, wall-socket energy; the best-grounded like-for-like number
1.4 pJ per H100 multiply-add vs 0.13 pJ per synaptic event (this page)10each machine's power divided by its own native operation
Sterling & Laughlin: a brain with a laptop's volume and power vs a supercomputer (Introduction; ch. 15)10⁵space and energy for a whole machine against Deep Blue, not per operation; the book attributes it to staying low on the rate–cost curve, not to better devices
"An exaflop at 20 W" vs Frontier's 1.1 EFLOP/s at 21 MW10⁶a number from the top of the whole-brain-emulation range against a measured LINPACK; not the same operation
Distance from Landauer: cortical neuron 10⁸ (Levy & Calvert); 8-bit add 10⁷; H100 per HBM bit 10⁹–10¹⁰≈1per bit, biology and silicon sit at the same distance from the thermodynamic floor; neither is near it

Task-level estimates of the brain in FLOP/s span 10¹² to 10²⁸ in the literature and are all one multiplication: synapses × rate × FLOPs per synaptic event, with the last factor free between 1 and 1,000. Carlsmith's mechanistic range, 10¹³–10¹⁷ FLOP/s, is the defensible band; at 20 W that is 10¹²–10¹⁶ FLOP per joule against the H100's 1.4×10¹² FLOP per joule at FP16.

Sources

H100. NVIDIA, H100 Tensor Core GPU Architecture whitepaper v1.02 (2022): 80 B transistors, 814 mm², TSMC 4N, 132 SMs, 256 KB RF/SM, 50 MB L2, NVLink 900 GB/s. NVIDIA H100 product page (2024): 989.5 dense FP16 / 1,979 dense FP8 TFLOPS, 67 FP32, 3.35 TB/s, 700 W. Luo et al., Benchmarking and Dissecting the Nvidia Hopper GPU Architecture, arXiv:2402.13499 (2024): 4,472 B/clk L2, 127.9 B/clk/SM shared memory. arXiv:2604.11391 (2026): 1,980 MHz boost, ~120 W memory power. NVIDIA Research, 0.190 pJ/bit 25.2 Gb/s die-to-die link in 5 nm (2023): 0.75 V. Keller, Venkatesan, Dally et al., JSSC 58(4) (2023): 95.6 TOPS/W INT4 at 5 nm. NVIDIA, Grace Hopper Superchip Architecture In-Depth (2023): NVLink-C2C 1.3 pJ/bit. DGX H100 datasheet: 10.2 kW, 130 kg.

Circuits. Horowitz, Computing's energy problem (and what we can do about it), ISSCC 2014, pp. 10–14, and its slide deck. Ho, Mai & Horowitz, The future of wires, Proc. IEEE 89(4) (2001): 55–80 ps/mm, 1.5–2 fF/µm. Weste & Harris, CMOS VLSI Design, lecture 7: activity 0.1 logic, 0.02 memory. Keckler et al., GPUs and the future of parallel computing, IEEE Micro (2011): 240 fJ/bit-mm per transition at 40 nm. Dally, Energy efficiency and AI hardware, Stanford AHA retreat (2023) and Hot Chips 2023: 100 fJ/bit-mm, 1 fJ/bit add, 30 pJ instruction overhead. O'Connor et al., Fine-grained DRAM, MICRO-50 (2017): 3.97 pJ/bit HBM2. Dreslinski et al., Near-threshold computing, Proc. IEEE 98(2) (2010). Zhai et al., DAC 2004. Restle et al., JSSC 36(5) (2001). Caputa, Linköping dissertation (2006). Landauer, IBM J. Res. Dev. 5(3) (1961).

Brain anatomy. Azevedo et al., J Comp Neurol 513 (2009): 86.1 B neurons, 16 B cortical. Pakkenberg & Gundersen (1997), Pakkenberg et al., Exp Gerontol 38 (2003), Tang et al., Synapse 41 (2001): 1.5–1.64×10¹⁴ neocortical synapses. Alonso-Nanclares et al., PNAS 105 (2008): synapse density. Marner et al., J Comp Neurol 462 (2003): 176,000 km myelinated fibre. Laughlin & Sejnowski, Science 301 (2003): 4 km axon/mm³, ~1 Tbit/s cortex, 1–16% active. Aboitiz et al. (1992), Tomasch (1954): callosal fibre counts. Caminiti et al., PNAS 106 (2009): callosal velocities and delays. Gentet et al., Biophys J 79 (2000); Eyal et al., eLife 5 (2016): membrane capacitance. Curcio & Allen, J Comp Neurol 300 (1990).

Brain energy. Attwell & Laughlin, JCBFM 21 (2001): 3.84×10⁸ ATP per action potential, 1.64×10⁵ per vesicle, 47/34/13/3% budget. Howarth, Gleeson & Attwell, JCBFM 32 (2012): 21/50/20% revised budget. Carter & Bean, Neuron 64 (2009): 1.24 sodium overload. Lennie, Curr Biol 13 (2003): 2.4×10⁹ ATP per spike, 0.16 Hz sustainable, 100 mm axon per neuron. Levy & Calvert, PNAS 118 (2021): 17 W glucose, 6.19 ATP-W, 3.5 W communication vs 0.10 W computation, 5.1 bit/s per neuron, 10⁸ from the physics limit. Harris & Attwell, J Neurosci 32 (2012): per-length axonal cost, myelin break-even 12.4 Hz. Perge et al., J Neurosci 32 (2012) and 29 (2009). Herculano-Houzel, PNAS 109 (2012) and PLoS ONE 6 (2011). Madsen et al., JCBFM 13 (1993). Georgiev et al., Brain Res 1748 (2020). Balasubramanian, PNAS 118 (2021).

Design principles. Sterling & Laughlin, Principles of Neural Design, MIT Press (2015): ATP = 25 kT and a G-protein bit for 25 kT (pp. 122–123); digital multiply steps (p. 131); pump = 60% of brain energy (p. 138); 1 ms of Na⁺ current = 2,000 ATP (p. 142); pump limit and a cortical mean below 10 Hz (pp. 152–154); 10⁹ kT per bit in a neuron (p. 157); analogue vs spikes, finest axons (pp. 174–175); channel-noise floor (p. 185); fly photoreceptor cost and capacity (pp. 224–228); single synapse vs array (p. 245); lamina redundancy (pp. 249–252); ganglion-cell bits per spike (p. 320); 2 Hz = resting cost (p. 342); 3/5 wire and 45,000× (pp. 367, 383–384); rate doubling (p. 393); spiking ≈20% of cortex (p. 395); storage (pp. 404, 411, 430); summary and the ten principles (pp. 433–446); 1.5×10⁶ km of axon (ch. 13, note 1). Balasubramanian & Sterling, J Physiol 587 (2009). Niven, Anderson & Laughlin, PLoS Biol 5 (2007). Howarth, Gleeson & Attwell (2012) for the 2 Hz pivot.

Information. Rieke, Warland, de Ruyter van Steveninck & Bialek, Spikes (1997); Strong et al., PRL 80 (1998): 90 bit/s. Koch et al., Curr Biol 16 (2006): 875 kbit/s guinea-pig retina, 1–3.5 bit/spike. Laughlin, de Ruyter van Steveninck & Anderson, Nat Neurosci 1 (1998): 10⁴ ATP per bit at a synapse. Zheng & Meister, Neuron 113 (2025): 10 bit/s; comment in Nat Neurosci (2025). Zimmermann (1989) for the 10⁹ bit/s sensory figure. Alle & Geiger, Science 311 (2006).

Comparisons. Naveen Rao, All-In Summit talk on the All-In Podcast channel, posted 2026-09-21, 08:25–08:57; Sequoia AI Ascent talk, 2026-05-06. Sarpeshkar, Neural Computation 10 (1998). Merkle, Foresight Update 6 (1989). Moravec (1998). Sandberg & Bostrom, Whole Brain Emulation: A Roadmap (2008). Sandberg, Energetics of the brain and AI, arXiv:1602.04019 (2016). Carlsmith, How much computational power does it take to match the human brain? (Open Philanthropy, 2020). Mehonic & Kenyon, Nature 604 (2022). Senk et al., Constructive community race, arXiv:2505.21185 (2025). AI Impacts, Brain performance in TEPS / FLOPS. TOP500 (Frontier). Mead, Proc IEEE 78 (1990), not verified against the primary text.

Sutro. Naveen Rao's energy heuristics, checked (sutro-problems, Sept 2026) for H7–H13; A100 numbers to know for the Dally and Horowitz constants and the A100 measurements.