Suncatcher: AI Compute in Orbit

Google is about to put TPUs into orbit as the first step toward solar-powered machine-learning clusters in space. The launch post is short. The engineering argument is in the paper behind it. This course lays out that argument: why the satellites must fly a few hundred metres apart, what the radiation tests showed, and how launch prices would have to fall. It separates what is demonstrated from what is projected, and names what the comparison with data centres leaves out.

6 modules
2 interactive tools
12 reasoning questions
~50 min

Built on Towards a future space-based, highly scalable AI infrastructure system design (Google, arXiv v2, Jun 2026) and Behind Project Suncatcher (24 Sep 2026)

How to read this course

The authors’ own claim is modest and precise: the core concepts are “not precluded by fundamental physics or insurmountable economic barriers.” That is a feasibility argument, not a business case. The paper says of its cost comparison, “The below does not constitute a full economic analysis.” The course holds them to that framing in both directions. It does not dismiss a feasibility result for lacking economics, and it does not read a feasibility result as a forecast.

Short on time? Do Module 2 (why they fly so close) and Module 5 (load-bearing versus speculative).

Course Modules

  1. The claim, and what launches nowStart here
  2. Why the satellites fly 100 metres apartInterconnect
  3. Radiation: what the proton beam showedHardware
  4. The launch-cost betInteractive
  5. Load-bearing versus speculativeBoth sides
  6. What the prototype can and cannot settleInteractive
1

The claim, and what launches now

The energy premise, the architecture in one paragraph, and how the plan has moved
By the end of this module you will
  • State the premise and the architecture in a few sentences
  • Distinguish the 2025 research proposal from the 2026 milestone

The premise

If AI is a general-purpose technology, demand for compute and energy will keep growing. The Sun emits more than 100 trillion times humanity’s total electricity production. The paper argues that “at some point in the future, the best way to power AI will likely thus be to more directly tap into that power.” The practical hook is orbit. In the right low orbit, satellites see near-constant sunlight, “generating up to eight times more solar power than on Earth,” with little need for batteries.

The architecture

No giant orbital data centre assembled in space. Instead a swarm of small satellites, each with solar arrays, TPUs and optical links. They fly in a dawn–dusk sun-synchronous low Earth orbit, which keeps them over the day–night line and in near-continuous sun. They fly close together so that laser links between them can match the bandwidth of the fibre inside a terrestrial TPU pod. Each satellite would eventually carry “dozens of TPU chips.”

How the plan has moved

WhenWhat was said
Nov 2025 (research blog + paper)A learning mission with Planet, “slated to launch two prototype satellites by early 2027.”
Sep 2026 (launch post)A first satellite on SpaceX’s Transporter-18 rideshare, built with Planet, to test TPUs in spaceflight, radiation and vacuum cooling. The laser links wait for two satellites in 2027.

The sequencing is sensible: prove the chip survives before proving the chips can talk. But it means the defining capability of the design, high-bandwidth links between close-flying satellites, is not what launches now.

Takeaways
  • Premise: abundant solar power in the right orbit
  • Design: a close-flying swarm of small TPU satellites, not a monolith
  • 2026 tests survival and cooling; the links come in 2027
2

Why the satellites fly 100 metres apart

The bandwidth gap, the inverse-square law, and the formation that closes it
By the end of this module you will
  • Explain the bandwidth gap between space lasers and a TPU pod
  • Explain why flying close closes it, and what that costs in orbital mechanics

The gap

Large ML jobs are spread across many chips that exchange data constantly. Inside a terrestrial TPU pod, chips talk over optical links at hundreds of gigabits per second per chip. Commercial satellite-to-satellite optical links run at roughly 1–100 Gbps. The paper sets a target of about 10 Tbps per link, one to four orders of magnitude beyond what space lasers do today. The 2026 post puts it plainly: existing systems are “optimized for low bandwidth across large distances,” while these need “very high bandwidth over extremely short distances.”

The physics that closes it

Received optical power falls with the square of distance. Long-range space links make do with about a microwatt at the receiver, which limits them to a few simple channels. The commercial dense wavelength-division multiplexing (DWDM) transceivers used in data centres, which put many wavelengths on one fibre, need hundreds of microwatts. Fly the satellites close enough and there is enough light for data-centre optics to work in free space.

  • 24 wavelengths at data-centre rates give about 9.6 Tbps on one link; tighter spacing could give 12.8 Tbps.
  • At even shorter range, several parallel beams (spatial multiplexing) fit in the same aperture, multiplying bandwidth again.
  • A bench demonstrator built from off-the-shelf parts achieved 800 Gbps each way (1.6 Tbps total) over a short free-space path.

The formation

The paper’s illustrative cluster is 81 satellites in one plane, at a mean altitude of 650 km, with a cluster radius of 1 km. Neighbouring satellites oscillate between roughly 100 and 200 m apart over each orbit. That is far closer than any existing constellation flies. Modelling with the Hill–Clohessy–Wiltshire equations plus a differentiable model in JAX suggests the cluster’s shape stays predictable. The main perturbation, Earth’s oblateness, can largely be designed around, so formation flight should need only “modest” manoeuvring beyond ordinary station-keeping.

The trade the design makes

Closeness buys bandwidth with commodity optics, but it spends safety margin. A hundred satellites a hundred metres apart is a new operating regime. Precision navigation, collision avoidance and the behaviour of a cluster when one satellite fails all become central. The paper models the orbits. It does not yet demonstrate close-formation operations, and that is what 2027 is for.

Takeaways
  • The design needs ~10 Tbps links; space lasers today do 1–100 Gbps
  • Proximity plus the inverse-square law lets commodity data-centre optics close the gap
  • The price is a formation 100–200 m apart, modelled but not yet flown
3

Radiation: what the proton beam showed

Survival looks solved; silent errors in training do not
By the end of this module you will
  • Read the radiation results against the mission requirement
  • Explain why the same error rate is fine for inference and open for training

Google tested Trillium (v6e) TPUs in a 67 MeV proton beam at UC Davis’s Crocker Nuclear Laboratory while they ran ML workloads. It reports this as the first published radiation test of such a device. With typical shielding, the expected dose in orbit is about 150 rad(Si) a year, so a five-year mission needs survival to about 750 rad(Si).

EffectResultAgainst the requirement
Total dose: long-term damageHBM memory showed irregularities from 2 krad(Si). Everything else ran correctly to 15 krad(Si). No hard failures.HBM margin ~3×; logic margin 20×
Silent data corruption: single-particle bit-flips in logic and SRAMAbout 1 event per 17 radOn the order of 1 per 3 million inferences, at 1 inference a second
Functional interrupts: host crashes and rebootsAbout 1 per 450 rad (CPU) and 1 per 400 rad (RAM)Less than one host reset per system per year

The authors’ own judgement on silent corruption: the rate “is likely acceptable for inference, the impact of SEEs on training jobs, and the efficacy of system-level mitigations, requires further study.”

Why training is different

An inference is short and independent. A rare wrong answer affects one request. A training run is one enormous computation spread across thousands of chips for weeks. An undetected bit-flip can propagate into the weights and quietly damage the whole run. Per chip, 150 rad a year at one event per 17 rad is about nine silent errors a year. Multiply by the number of chips in a cluster and the count is no longer rare. Silent data corruption is already a known problem on the ground. Space raises the rate, and so raises the importance of the detection and correction the paper says still need study.

Takeaways
  • Survival: no hard failures to 15 krad against a 750 rad need; HBM is the weak link
  • Silent errors: fine for inference, open for training
  • Per-chip rates become daily events at cluster scale
4

The launch-cost bet

A learning curve, a fleet of Starships, and a comparison that is deliberately partial
By the end of this module you will
  • Explain the learning-curve projection and what has to happen for it to hold
  • Reproduce the paper’s “launched power price” and its comparison with data-centre power

The learning curve

SpaceX’s price history, from Falcon 1 to Falcon Heavy, fits a learning rate of about 20%: the price per kilogram falls about 20% for every doubling of cumulative mass launched. If that continues, which “would require ~180 Starship launches/year,” launch to low Earth orbit could fall below $200/kg by about 2035. With about 72% less mass launched, the same curve gives about $300/kg. A bottom-up look at Starship’s published specifications points the same way: roughly $60/kg to SpaceX with 10× reuse, under $15/kg with 100× reuse, and a propellant floor near $8/kg. Today’s reference is $3,600/kg on a reusable Falcon 9.

The launched power price

To compare with a data centre, the paper asks what it costs to put a kilowatt of solar power in orbit, per year of life. For a Starlink v2 mini (575 kg, about 28 kW, 5-year life): at $3,600/kg that is about $14,700 per kW per year; at $200/kg about $810. Across other satellite designs, the $200/kg range is $810–7,500. US data-centre power costs about $570–3,000 per kW per year ($0.06–0.25/kWh, PUE 1.09–1.4). So at $200/kg, launched power could be “roughly comparable” to terrestrial energy spend.

Tool · Launched power versus grid power

Reproduce, then stress
Launch price$200/kg
Satellite mass575 kg
Power per satellite28 kW
Useful life5 yr
Grid price$0.10/kWh
Data-centre PUE1.20
What this shows

Launched power price = launch price × mass ÷ kW ÷ life, as in the paper. Grid power = price × 8,760 h × PUE. Like the paper, this excludes chips, satellite manufacture and data-centre buildings.

Takeaways
  • ~20% learning rate → <$200/kg by ~2035, if ~180 Starships a year fly
  • At $200/kg, launched power is within the range of terrestrial power spend
  • The comparison is power only, by design; space-only hardware is outside it
5

Load-bearing versus speculative

Which parts of the argument are demonstrated, which are modelled, and what is not yet addressed
By the end of this module you will
  • Sort each claim by how much evidence sits behind it
  • State the strongest objections, clearly marked as objections
ClaimStatusWhy
Up to 8× more solar energy in dawn–dusk orbitPhysicsOrbital geometry and insolation; not in dispute
TPUs survive a five-year doseTestedProton-beam tests with margin; in-orbit data from the 2026 mission
Terabit free-space optics with commodity partsBench1.6 Tbps on the bench; not between moving spacecraft
Stable 100–200 m formation of 81 satellitesModelledOrbital-dynamics modelling; not flown
Silent errors acceptable for trainingOpenAuthors say it “requires further study”
Launch <$200/kg by ~2035ProjectionExtrapolated learning curve; depends on one vendor’s volume
Cost-competitive with terrestrial computeNot claimedPower-only comparison; “not a full economic analysis”
Objection 1

Heat, not power, may be the binding constraint

A communications satellite spends much of its power on radio and sends some of it away as signal. A compute satellite turns nearly all of its power into heat in a small area, and in vacuum “you can only diffuse heat via radiators.” Radiator area grows with power, which pushes up kg per kW. The launched-power arithmetic borrows a Starlink’s mass-to-power ratio, which may flatter a compute design.

The best reply: the paper calls thermal management a “critical optimization challenge,” and cooling is one of the three things the 2026 satellite tests. The tool in Module 4 shows how much heavier the design can get before the comparison fails.
Objection 2

You cannot swap a failed chip in orbit

“Currently, failed TPUs are manually replaced by technicians,” the authors note, which is “obviously impracticable in space.” Their answer is redundancy. Spare capacity is paid for at launch prices and carried for the whole mission.

The best reply: redundancy is standard spacecraft practice, and work on fault-tolerant training reduces how much is needed. It is a cost line, not a blocker. It belongs in the economics the paper deliberately left out.
Objection 3

Five years is long for an accelerator

The launched power price amortises over five years. On the ground, operators replace accelerators when newer chips are enough better per watt. In orbit you cannot refresh chips without launching new satellites, so obsolescence either shortens useful life, which raises the price per kW-year, or strands old compute in orbit.

The best reply: a modular swarm can be refreshed satellite by satellite, and at low launch prices replacement is routine. The argument depends on launch price again, which makes the launch projection even more load-bearing.
Objection 4

The data still has to come down

Useful work needs inputs up and results down. The best demonstrated optical ground link is NASA’s TBIRD at 200 Gbps (2023). The paper names atmospheric turbulence and beam tracking as challenges. This favours workloads that are compute-heavy and data-light, like some training or batch inference, over interactive serving.

The best reply: the paper treats radio as enough for the pilot and optical downlinks as later work. The honest reading is that orbital compute would serve particular workloads, not replace data centres.
Takeaways
  • Physics and radiation survival are solid; links and formation are bench-tested or modelled; training errors and economics are open
  • Heat, repair and obsolescence are real costs outside the power comparison
  • Nearly every reply leans on cheap launch, which makes it the load-bearing projection
6

What the prototype can and cannot settle

Reading a first mission for what it will actually tell us
By the end of this module you will
  • Say which open questions the 2026 and 2027 missions can retire
  • Know what evidence would genuinely change the outlook

The post frames the first mission modestly: “This first launch is about seeing what works, identifying points of failure, and applying those findings to future missions.” Before launch, the hardware already survived vibration testing on all three axes, with loads of up to 10 g for the vehicle and 50–100 g for components like the chips. “Tests like this rarely go as planned, so we were pleasantly surprised that the hardware held up.”

Tool · Which mission settles it?

Commit, then see why0 of 6 sorted

For each open question, decide whether the 2026 single-satellite mission can largely answer it, whether it waits for the 2027 two-satellite mission, or whether neither can.

Do TPUs and their heat-pipe and radiator cooling keep working in vacuum and real orbital radiation?
2026. This is what the first satellite is for. It gives in-orbit data to check against the beam tests and the thermal-vacuum chamber.
Can two satellites hold a high-bandwidth laser link while both are moving?
2027. It needs two satellites. Google says it will test this in 2027.
Will launch fall below $200/kg by the mid-2030s?
Neither. It depends on launch-industry volume and reuse, not on anything Suncatcher flies.
Can an 81-satellite cluster fly safely with 100–200 m spacing?
Neither, fully. Two satellites show relative navigation and pointing. Many-body formation behaviour, failure handling and collision risk at cluster scale need a much larger demonstration.
Is the silent-error rate acceptable for large training runs?
Neither. A few chips can measure the per-chip rate in orbit. Whether training tolerates it at scale is a systems question: detection, correction and fault-tolerant training across thousands of chips.
Does the real orbital radiation dose match the ~150 rad(Si)/year estimate?
2026, at least in first form. On-orbit dosimetry and error counts test the model that every radiation margin in Module 3 depends on.
What would genuinely change the outlook
  • Up: sustained Starship cadence with high reuse; a two-satellite link near terabit rates; in-orbit error rates at or below the beam-test prediction.
  • Down: radiator mass per kW far above communications satellites; training runs that cannot tolerate orbital error rates without heavy redundancy; launch cadence stalling well short of the learning curve.
Takeaways
  • 2026 retires survival and cooling risk; 2027 tests links
  • Launch cost, formation at scale and training reliability are beyond either mission
  • Watch Starship cadence as closely as Suncatcher itself