Showing posts with label earth-observation. Show all posts
Showing posts with label earth-observation. Show all posts

Thursday, June 25, 2026

EO Platforms, Science Platforms, and the Apartment Between Binder and Earth Engine

When you think of running code in the browser, you probably think of Binder or Google Colab. If you are at an AWS house, you think of Sagemaker.

That is the visible end of the digital science platform story: a notebook appears in a tab, Python imports something, a chart lands on the screen, and for a moment the whole research infrastructure problem looks solved. If you are fan of a different paradigm, Python can be replaced with R, or a programming language replaced by a chat session which has its side effects as Science/Engineering plans and some code.

Then the real workflow arrives. The data is 100 TB of imagery from a High-Altitude Balloon and it is sitting at the HPC; you will have to wait in the queue for the job to submit, but the proposal deadline is tomorrow and you really need preliminary exploration graphics to illustrate your story, since humans are inherently visual creatures. Do you fake it till you make it using Nano Banana, or do you commit to a fraction of the actual work you are applying for the funding to do in full?

The HPC credentials are institutional. The analysis needs Dask workers near cloud object storage. It will take two hours for Globus to haul the data over to the nearest AWS Region. The output, when the grant is completed, should become a tile service, a STAC collection, a policy brief, a model run, and a thing someone can repeat six months later when the next sensor collection lands after that much needed CMOS and Lens upgrade. For now you really need that plot showing how a thermal mosaic align with the visible spectrum satellite imagery.

At that point the browser notebook is the doorway, not the platform. The doorway opens onto a howling void of needs and unfilled expectations.

The apartment, the hotel room, and the empty block

The continuum feels a bit like the Django versus Flask argument, except measured in satellite scenes, climate arrays, grant budgets, and institutional risk. If you don’t know what Django or Flask is, perhaps opinionated vs minimalistic are more understandable.

A managed science namespace is the Flask or FastAPI option. You get a well-tested core, working access controls, shared storage, a scheduler, a route to scalable compute, and enough freedom to design your own universe. It is like buying an empty apartment in a good part of town. The walls are sound. The plumbing works. The rooms are yours, you will not get robbed if you step outside.

A highly managed platform is closer to a hotel room at the Ritz. Everything is convenient, polished, and serviced. You get a desk, a bed, a view, and a button to press when something breaks. You probably cannot knock down a wall.

Then there is the empty block of land: Kubernetes, object storage, IAM, container registries, workflow engines, data catalogs, and a team with enough scars to make them all behave. Maximum control, maximum blast radius.

The interesting work sits between those extremes. Digital science platforms are not one product category. They are a gradient of convenience, control, reproducibility, and institutional memory.

The first split: interactive science versus production exploitation

Earth observation platforms and broader science platforms grew from different pressures.

Science platforms usually start with researchers. They need notebooks, environments, collaboration spaces, package installation, teaching material, shared compute, and a path from experiment to publication. The center of gravity is interactive work: Jupyter, RStudio, Dask, documentation, executable papers, and a community that can teach the next person how to reproduce the result.

EO exploitation platforms usually start with data gravity. They need catalogs, tiling, pixel access, area-of-interest filtering, atmospheric correction, process graphs, long-running jobs, and application packaging. The center of gravity is operational work: find the scenes, run the process near the archive, publish the output, and make the result portable across platforms.

In practice those two worlds are merging. The Pangeo stack wants cloud-native scientific computing. EOEPCA wants interoperable EO exploitation building blocks. Digital Earth platforms want curated national-scale data cubes. NASA-style platforms such as MAAP and VEDA want science teams to move from discovery to analysis to communication without rebuilding the pipes each time.

Choosing a winner is not the problem. Knowing which layer you are buying is.

A rough map of the ecosystem

The ecosystem makes more sense sorted by the job each platform does rather than by brand.

Layer Main job Examples What you gain What you give up
Ephemeral notebooks Launch a runnable environment from code Binder, repo2docker-backed demos Low friction, great teaching surface Weak persistence, weak governance, limited scale
Hosted notebooks Put compute in a managed browser workspace Google Colab, Kaggle notebooks, SageMaker Studio Lab-style services Fast start, familiar interface, low admin load Runtime limits, cloudy provenance, platform-specific behavior
Community hubs Give a research community a managed shared workspace 2i2c hubs, JupyterHub deployments, Pangeo hubs Shared identity, shared environments, scalable Dask, community ownership Still needs platform stewardship and cost discipline
Open science practice layers Teach teams how to work reproducibly OpenScapes, Project Pythia, Jupyter Book, MyST Social machinery for reproducible science Not a compute platform by itself
Cloud-native geoscience stacks Make large arrays workable near object storage Pangeo, Xarray, Dask, Zarr, Kerchunk, Intake Open, composable scientific computing You assemble and operate the stack
Data cube and EO libraries Organize analysis-ready EO collections Open Data Cube, xcube, stackstac, odc-stac EO-shaped data abstractions Needs catalog, storage, and operations around it
Managed planetary platforms Bring data catalog and compute together Google Earth Engine, Microsoft Planetary Computer Convenient data proximity and APIs Platform dependence and varying escape hatches
Standards-based EO exploitation Package processing for portability openEO, OGC API - Processes, EOEPCA, CWL-style application packages Repeatable jobs across back ends More ceremony and specification work
Science mission platforms Align infrastructure to a mission community NASA MAAP, NASA VEDA, AquaWatch, ESA EarthCODE-style efforts Domain-specific data, workflows, and communication Mission boundaries shape what is easy
Institutional digital earths Curate national or sectoral data products Digital Earth Australia, Africa, Pacific, and related platforms Trusted local data products, policy relevance Sustained funding and data governance required

That table is too tidy, of course. Real platforms straddle rows. A Pangeo hub can be a community hub, a cloud-native geoscience stack, and the working surface for a Digital Earth program. Earth Engine is both a managed planetary platform and a workflow language. EOEPCA is both architecture and open-source building blocks. A Digital Earth platform can host notebooks, APIs, dashboards, STAC catalogs, and products built from Open Data Cube.

The taxonomy still helps because “we need a science platform” can mean a teaching hub, a production EO processing system, a mission workbench, a public communication portal, or a Kubernetes tenancy model for research groups. Those are different purchases.

Pangeo is a community before it is a stack

Pangeo is often described through its software: Xarray for labelled N-dimensional arrays, Dask for parallel compute, Zarr for chunked cloud-friendly storage, Jupyter for interactive work, Kerchunk for virtualizing old file formats, Intake for catalogs, hvPlot for quick visualization, XPublish for serving datasets, and newer bounded-memory or serverless experiments such as Cubed and Xarray-Beam.

What goes into the requirements / constraints file of a notebooks repo is all and good, but like all code its value really depends on the people who read it and run it, Anne Frank’s diary is no good, if it is never found and read. The written word of any form be is a notebook or a painful diary, reaches its true potential when it transmits the ideas contained there in to the minds of thousands of others. Pangeo describes itself as a community for open, reproducible, scalable geoscience. Project Pythia is its education working group and a training hub for the geoscientific Python community. The infrastructure, documentation, showcase talks, working groups, cookbooks, and foundation tutorials are part of the platform. Project Pythia can be found, it has a lot of words, but no concrete infrastructure to run on, no printing press to replicate the proposed patterns and keep publishing them across planning cycles.

Many institutional platform programs miss this. You can install JupyterHub and Dask Gateway in a week. You cannot apt install a culture of reproducible workflows, documented environments, data access patterns, and examples that match the local science questions.

2i2c and OpenScapes occupy this gap. 2i2c turns the hub into managed open infrastructure with access control, custom environments, scalable compute, and a “Right to Replicate” philosophy, which is a useful antidote to the managed-service trap. OpenScapes works on the human operating system: habits, onboarding, documentation, and shared practice. The difference shows up when one champion leaves and the platform either survives or becomes a quiet monthly cloud bill, that everyone questions every quarter and decides to shutdown in a much “killed by Google” fashion. It have shut down a service or two in my time as well.

Earth Engine is the polished hotel room

Google Earth Engine remains the reference point for convenient planetary-scale EO analysis. It combines a multi-petabyte public catalog, server-side geospatial computation, JavaScript and Python APIs, and a browser code editor that made remote sensing feel weirdly immediate.

The trade-off is the same one that makes it attractive. You get the hotel room. You do not own the building.

For many users that is exactly right. If the job is land-cover change detection, trend mapping, rapid prototyping, teaching, or humanitarian analysis, Earth Engine collapses a brutal infrastructure problem into an API, and it has unlocked a large amount of science and public-interest work.

A national agency or regulated enterprise eventually asks harder questions. Can we run this against restricted internal data? Can we reproduce the process outside the platform? Can the cost model survive operational use? Can the workflow be audited? Can the output become an internal service rather than a script in someone else’s editor? What if Google does not like the economics and GEE gets “Killed by Google” ?

Those questions mark the boundary between convenience and control.

Planetary Computer, STAC, and the cloud-native middle

Microsoft Planetary Computer pushed a different shape: open catalogs, STAC metadata, cloud-hosted environmental datasets, Hub-style analysis environments, and APIs that fit the broader Python geospatial ecosystem. It did not only put data in the cloud; it made the catalog and assets legible to ordinary tools.

That middle ground is where STAC, Cloud Optimized GeoTIFF, Zarr, GeoParquet, and object storage conventions earn their keep. STAC calls itself a common language for geospatial information, with a core JSON structure for describing and cataloguing spatiotemporal assets. That sounds dull until you have inherited five satellite archives, three naming conventions, and one research assistant who left for a postdoc in Bremen.

These are the unglamorous freight standards of cloud-native geospatial work. Once data is published in these shapes, platforms compete on experience and operations instead of trapping users at the file boundary.

The same applies inside Digital Earth programs. Open Data Cube, odc-stac, xcube, stackstac, and Pangeo tooling become more powerful when the data model is boring and inspectable. Boring is good. Boring is how a wet-season flood product, a crop monitoring workflow, and a methane detection experiment can reuse the same infrastructure without pretending they are the same science.

openEO and EOEPCA are the uncomfortable middle

At the other end from free-form notebooks are workflow and process standards.

openEO gives users a unified API for Earth observation cloud back ends. It is a process graph world: define operations, submit work, let the back end execute near the data, and avoid rewriting the same analysis for every platform. In 2026 openEO became an OGC Community Standard, which is a useful signal that the idea has moved beyond a single project vocabulary.

OGC API - Processes handles a related need: wrap computational tasks as executable processes that a server can offer through a JSON-over-HTTP Web API. It is the unromantic contract that lets processing become a service instead of a notebook cell someone has to rerun by memory.

EOEPCA then packages a broader exploitation platform architecture around this world: reusable building blocks, open standards, federation between EO cloud platforms, application packaging, identity, data access, processing, and publication. Its mission is practical: enhance interoperability between cloud-based EO platforms and unify the fragmented cloud ecosystem for ground segment, EO science, R&D, and applications.

This middle is uncomfortable because it asks scientists to describe workflows with more ceremony than a notebook while asking infrastructure teams to expose more flexibility than a fixed portal. That ceremony earns its keep only when portability, audit, repeatability, and multi-platform execution are real requirements.

MAAP, VEDA, and AquaWatch show mission-shaped patterns

NASA’s Multi-Mission Algorithm and Analysis Platform (MAAP), Visualization, Exploration, and Data Analysis (VEDA), and AquaWatch point at another category: mission-shaped science platforms.

MAAP is built around collaborative algorithm development and analysis for biomass, carbon, and related mission data. It combines data, algorithms, and cloud computation for Earth science work at scale, including GEDI, ICESat-2, BIOMASS, NISAR, and related mission products. Its centre is not a generic notebook. It is a community of practice around specific data products, algorithms, and validation work.

VEDA is more communication and exploration oriented: science data, dashboards, stories, APIs, and applications that help users move from dataset to visible public insight. Its public dashboard talks about cloud-enabled analysis without writing code or downloading data, plus stories for science communication. In the geospatial agent stack draft I have been calling this the presentation and evidence layer: not just compute, but a way to package outputs for humans without losing provenance.

AquaWatch is the mission model closer to my own recent work: water quality intelligence rather than a generic EO workbench. Its platform shape is dictated by the problem. Satellite products, in-situ observations, calibration and validation workflows, data services, state-government users, and operational water decisions all have to sit in one fabric. Nobody wants an elegant notebook if the harmful algal bloom window has already closed and the river manager is still waiting for a plot.

Mission platforms should not be embarrassed by their specificity. A platform that knows its science domain can make strong defaults: data already mounted, examples already relevant, units already sensible, common plots already nearby, and metadata that matches the decisions people actually make.

Digital Earth platforms are institutional memory

Digital Earth platforms are a cousin of science platforms, but with more institutional weight.

Digital Earth Australia, Digital Earth Africa, Digital Earth Pacific, and related efforts are not merely places to run notebooks. They are long-running data product factories, national or regional knowledge systems, and trust infrastructure. They turn raw satellite archives into analysis-ready data and derived products that policy, environment, agriculture, water, and disaster teams can use.

The local flavours carry the point. Digital Earth Australia talks about trusted imagery and freely available analysis-ready datasets for researchers, land managers, agriculture, and emergency management. Digital Earth Africa frames the work as decision-ready products for sustainable growth across social, environmental, and economic challenges. Digital Earth Pacific is explicitly about decision-making for Pacific peoples at EO scale, with regional products, dashboards, an analytical hub, data, community, and governance in the same frame. This is not just a different skin on the same JupyterHub.

If you need a global flavour you can check out EASI from CSIRO under the AquaWatch program delivery, it flips the regional data cube approach to provide a global datacube from STAC using cloud native primitives. Almost a build your own Google Earth Engine approach, with associated DevOps and Platform Engineering burdens which you have to pay for to get.

That is a different job from a research notebook service. The value is not that one scientist can compute something, but that many people can rely on the same corrected, versioned, documented product line. Open Data Cube sits under many of these programs as an open-source way to manage analysis-ready data and has been used for continental-scale products such as Australian land-cover mapping from decades of Landsat imagery.

Here the apartment analogy changes. A Digital Earth platform is not an apartment you furnish. It is an apartment building with strata rules, fire exits, a maintenance schedule, and residents who will complain if the lift stops working, and I have been a super on such apartment buildings across 3 platforms now.

What must be portable?

Most platform debates get stuck on user interface preference. Notebooks versus portals. Python versus JavaScript. Managed service versus Kubernetes. Proprietary convenience versus open stack purity. The relevant debate is around what must be portable.

If only the result is needed, a polished managed platform may be perfect. If the method must move between agencies, clouds, and missions, then STAC, openEO, OGC API - Processes, CWL-style packaging, and open data formats become non-negotiable. If the community must outlive one grant, then docs, examples, governance, and training are platform features. If restricted data is involved, then identity, audit, and network controls are part of the science platform, not enterprise up-sell (don’t get me started on the SSO premium most SaaS platforms charge).

There are at least five kinds of portability hiding inside this one word:

  1. Data portability: can assets move or be read by standard tools?
  2. Workflow portability: can the analysis run on a different back end?
  3. Environment portability: can the software stack be rebuilt?
  4. Evidence portability: can another person inspect how the result was produced?
  5. Community portability: can another team inherit the practice without oral tradition?

The best platform choice depends on which of those you actually need.

Agents make the old platform boundary fuzzier

Agentic tooling is the next complication.

Once agents can call a STAC API, resolve a place name, submit an openEO job, monitor an Argo Workflow, query a Digital Atlas, open a notebook, and draft a VEDA-style story, the visible platform recedes behind the contracts underneath it.

Portals do not disappear. The durable value moves down into services, catalogs, workflow engines, identity boundaries, and audit trails, and the UI becomes one client among many.

This is where EO and enterprise AI rhyme. A user should be able to ask a normal question:

Show me whether wet-season surface water around this catchment is outside the last ten-year range, and give me the data sources and caveats.

The system behind that question might need a gazetteer, a STAC catalog, a data cube, Dask workers, an openEO process graph, a workflow engine, a notebook artifact, and a short written summary. Calling all of that a chatbot misses the point. It is a digital science platform with a language-shaped front door.

Science Platform advantage

There is a less glamorous question hiding under the agent story: who operates and pays for the tooling and compute fabric the agents call?

If a user asks the wet-season surface-water question and an agent fans out across a slew of tools, somebody owns every link in that chain. Somebody pays for object storage requests, scheduler time, network egress, GPU or CPU minutes, model tokens, human review, and the maintenance of the small apis that make the orchestration possible.

A serious science platform needs weight propagation in both directions. Backward propagation attributes each output to the tools, datasets, people, grants, teams, and compute fabric it consumed. Forward propagation estimates what the next run will cost before the button is pressed. Without both, the presenter of the agentic tooling with customer touch point gets all the glory and gets to keep all the revenue, while the sustaining ecosystem underneath withers on the vine and the agentic tool ends up being a toy in the long run.

This is where the phrase “Science Platform advantage” starts to be useful, in the same way people talk about quantum advantage. The claim is not that a platform is clever. The claim is that a connected, orchestrated chain of task-specific tooling, assembled on demand and cost-tracked properly, delivers more value than doing the same work any other way.

For the utilitarians in the room, that claim needs a benchmark. Pose a real problem. Solve it fully the old way first: portal clicks, downloads, hand-written scripts, email threads, queue waits, manual plots, confused provenance, and the final slide dropped into a proposal at 11:47 pm. Then build the agentic way and run the same class of problem a thousand times, so the cost of building the new fabric is amortized across many analyses rather than charged to the first heroic demo.

The trouble is that no one was keeping clean accounts while we did it the old way. We rarely know how much the portal-clicking cost once staff time, failed runs, duplicate downloads, local storage, interpretation drift, and meetings are counted. So we cannot honestly promise cheaper or faster without measuring both sides. We can be more certain about the energy bill: silicon currently delivers far less useful scientific reasoning per watt than the human brain fed on caffeine and croissants.

That does not weaken the platform argument. It makes the accounting part of the platform. The useful question is not “can an agent do this?” The useful question is whether the platform can tell us what it did, who paid, who benefited, and whether the next thousand runs beat the old way by enough to justify the extra silicon. Will the environmental solution we are building preserve the environment or become yet another nail in the anthropogenic coffin we are putting around our planet.

My current bias

I lean toward the apartment model.

Give science teams a managed namespace with identity, storage, standard catalogs, scalable compute, observability, and a small number of blessed paths to production. Keep the walls movable. Let the team bring Pangeo, Open Data Cube, xcube, DuckDB, GDAL, QGIS plugins, or specialist models where the science demands it. Make the durable contracts boring: STAC, COG, Zarr, GeoParquet, openEO, OGC APIs, container images, reproducible environments, and documented workflows.

Use the hotel room when speed beats ownership. Use the empty block only when the institution is ready to become a platform operator. Everything else is a furnishing choice.

EO platforms and science platforms are not rival tribes. They are different answers to the same infrastructure question: how much freedom do you need, how much responsibility can you carry, and what parts of the science must still make sense when the browser tab is closed?

Sources and threads to pull next

Saturday, March 7, 2026

Two Agents, One Codebase: An F1 Race Team Approach to Porting ACOLITE to Rust

In Formula 1, every team fields two drivers. Not as a backup plan – as a strategy. One driver pushes the pace, forcing rivals to respond. The other holds position, manages tyres, and covers the alternative strategy. They share telemetry, they share a garage, but they are running different races on the same track. The team wins when both cars score points, not when one driver tries to do everything.

Porting a scientific Python codebase to Rust feels remarkably similar. You need the aggressive driver – the one who charges into unfamiliar code and lays down fast laps of Rust implementation. And you need the calculating driver – the one who reads the data, watches for degradation, and calls out when the numerical precision is drifting. Two AI coding agents, paired like Norris and Piastri, sharing a codebase but operating on different parts of the problem.

The Starting Grid: Why Rust for ACOLITE?

ACOLITE is RBINS’ atmospheric correction toolkit for aquatic remote sensing. It handles everything from Landsat and Sentinel-2 to hyperspectral sensors like PACE OCI (286 bands) and PRISMA (239 bands). The Dark Spectrum Fitting (DSF) algorithm is elegant – image-based, no external atmospheric inputs – but in Python, processing a full PACE scene involves reading 291 NetCDF variables, interpolating multi-dimensional LUTs, and correcting each pixel’s reflectance through a chain of gas transmittance, Rayleigh scattering, and aerosol models. On a decent machine, this takes around 230 seconds.

The seed was planted at FOSS4G 2025 in Auckland when Leo Hardtke ran a tutorial on Earth Observation processing with Rust. It was plagued by Nix environment issues (as I noted in my conference write-up), but when the code ran, it was fast. Zero-cost abstractions and fearless concurrency are not just slogans at that point – they are wall-clock seconds you are not spending waiting for your atmospheric correction to finish.

I had also been watching Rob Woodcock’s acolite-mp branch, which tackled the same performance problem from within Python. His approach was clever: per-band parallelism with memory budgets tuned to cloud CPU-to-RAM ratios (2, 4, or 8 GiB per core), replacing NumPy’s interpolation with the multithreaded pyinterp, and carefully managing the GIL contention that Python’s threading model inflicts on you. He got Sentinel-2 from 791s down to 197s and Landsat from 312s to 99s on a 24-core i9 – roughly a 3-4x speedup.

But the GIL is still there. The memory model is still Python’s. And as Rob himself noted, “further performance improvements are possible but require more extensive changes to the file handling” and “there is a fair amount of GIL contention which limits threading being caused by some structural choices in the implementation.” At some point, you are fighting the language rather than the problem.

Rust sidesteps all of this. No GIL. No garbage collector. Rayon gives you data-parallel iterators that map across bands or tiles with work-stealing. Memory usage is deterministic and known at compile time – you can profile it statically before deploying, which is a sentence that makes no sense in Python-land but is table stakes in systems programming.

The Pit Crew: Two Agents via ACP

Here is where the teammate analogy really kicks in. In F1, a team with only one driver is not half a team – it is no team at all. You cannot run a split strategy with a single car. You cannot use one driver to hold up a rival while the other pulls a gap. The performance of the pair exceeds the sum of the individuals because they create options that a solo driver simply cannot.

Porting 40,000+ lines of scientific Python to Rust is the same. A single AI agent writing Rust will drift – the implementation slowly diverging from Python’s numerical behaviour until your reflectance values are off by just enough to be scientifically useless. You need the second driver to keep it honest.

The solution I landed on was a multi-agent orchestration harness using the Agent Client Protocol (ACP), a JSON-RPC 2.0 protocol over NDJSON stdio that lets coding agents communicate in a structured way:

Agent Role F1 Equivalent
Kiro Executor – writes Rust code, runs tests, reads files Lead driver – pushes the pace, sets fast laps
Copilot Proposer – reviews output, suggests next steps, cross-checks Python Second driver – covers the strategy, watches the gaps
Human Approver – filters proposals before dispatch Team principal – makes the call on when to pit

The workflow per sensor port looks like this:

  1. Human provides --task to the orchestrator (tools/agent_harness.py)
  2. Kiro receives the task via ACP session/prompt and starts writing code
  3. Kiro streams output via session/update chunks
  4. Output goes to Copilot for review against the Python source
  5. Copilot proposes ACTION: lines – “fix the gas transmittance interpolation order”, “the Rayleigh LUT needs pressure stacking”
  6. Human approves or rejects
  7. Approved actions go back to Kiro
  8. Repeat until regression tests pass or maximum cycles reached

This is not vibe coding. This is a two-car team running a split strategy.

Think about how McLaren or Red Bull operate. The lead driver qualifies on pole and sets the pace in clean air. The second driver starts on a different tyre compound, runs a longer first stint, and emerges from the pits into a different part of the field. They are solving complementary problems – one optimises for raw speed, the other for strategic coverage. Neither is redundant.

Kiro is the lead driver. It attacks the Rust implementation aggressively – writing loaders, porting DSF algorithms, wiring up rayon parallelism. It sets fast laps. It also occasionally bins it into the gravel trap by hallucinating a NumPy broadcasting rule that does not exist in ndarray.

Copilot is the second driver. It reads the Python source, cross-references the Rust output, and spots where the gap to parity is growing. “The gas transmittance interpolation order is wrong” is exactly the kind of radio call a second driver makes – not flashy, but it prevents a DNF.

The human is the team principal. You do not override the drivers on every corner, but you make the strategic calls: do we pit now and fix this RMSE regression, or do we push on and address it in the next stint? Is a 0.002 RMSE difference in Sentinel-2 reflectance acceptable? (It is – that is within float32 precision.) When do we switch from tiled DSF to fixed DSF mode for this sensor?

Together, they converge faster than either alone, for the same reason that two cars gathering tyre data in free practice gives the team more information than one car doing twice as many laps.

The Telemetry: Regression Tests Against Real Data

In F1, both drivers generate telemetry. The team overlays their data – braking points, throttle application, cornering speed – to find where one is faster and why. The overlay is the truth. Not the driver’s feeling, not the engineer’s simulation, but what the car actually did on the track.

Regression tests are our telemetry overlay. The Python ACOLITE output is Driver 1’s trace. The Rust output is Driver 2’s. We overlay them pixel-by-pixel, band-by-band, and look at the delta. When the traces diverge, something real has changed and we need to understand whether it is a genuine improvement or an error we need to correct.

There are currently 141 Python regression tests that compare Rust output against Python output pixel-by-pixel across real satellite scenes:

  • Landsat 8/9: 13 regression + 13 Rust-vs-Python + 7 benchmark tests
  • Sentinel-2 A/B: 19 regression + 15 Rust-vs-Python + 9 benchmark tests
  • PACE OCI: 17 regression + 14 Rust-vs-Python + 12 DSF comparison + 12 ROI + 10 full-scene tests

The tolerances are tight. Sentinel-2 achieves RMSE < 0.002 (physics-equivalent). Landsat gets RMSE < 0.02. PACE full-scene (1710 x 1272 pixels x 291 bands) hits mean RMSE of 0.004 with 100% of pixels within 0.05 of Python. Correlation coefficients are R > 0.999 across all sensors.

These are not toy tests on synthetic data. They run against actual L1 scenes downloaded from USGS and NASA. When the tests break, something real is wrong.

The Performance Gap: Where the Seconds Go

Sensor Scene Size Rust Python Speedup
Landsat 8 62M px x 7 bands 66s 180s 2.7x
Landsat 9 62M px x 7 bands 56s 180s 3.2x
Sentinel-2 A 30M px x 11 bands 52s 182s 3.5x
Sentinel-2 B 30M px x 11 bands 64s 173s 2.7x
PACE OCI (full) 1710 x 1272 x 291 bands 84s 230s 2.7x

The PACE result is particularly satisfying. The key optimisation was switching from 291 per-band NetCDF reads to 3 bulk detector reads, then applying rayon-parallel atmospheric correction across tiles. Load is 12 seconds, AC is 34 seconds, write is 35 seconds. That write phase for a 291-band hyperspectral cube goes to GeoZarr V3 with gzip compression – try doing that in a Python event loop without your memory allocator throwing a tantrum.

Energy Efficiency: The Fuel Strategy Nobody Talks About

Here is the part where I get philosophical – and where the F1 analogy turns from metaphor into mirror.

Formula 1 underwent a fuel efficiency revolution in 2014. The FIA introduced hybrid power units, capped fuel flow at 100 kg/hour (monitored 2,200 times per second), and forced teams to extract maximum performance from minimum fuel. The result was not slower cars – it was faster cars that used less. The 2026 regulations go further: fossil carbon is prohibited entirely, the MGU-K will deliver three times the electrical power (350kW vs today’s 120kW), producing up to 1,000 horsepower while burning sustainable fuel. Less fuel, more power. That is not a trade-off – it is an engineering constraint that drives innovation.

The same constraint applies to scientific computing, we just pretend it does not. Cloud computing bills are denominated in dollars, but the underlying unit is energy. Every CPU cycle your atmospheric correction burns is a watt drawn from a power grid somewhere. When you are processing continental-scale Sentinel-2 archives or the full PACE ocean colour mission, those watts add up. Python is the V10 era of scientific computing – glorious, unrestricted, and profligate with resources.

Rust is the hybrid power unit. Its advantage is not just speed – it is energy per unit of work. A 3x speedup roughly translates to using a third of the compute time, which means a third of the energy, a third of the carbon footprint, and a third of your AWS bill. The Rust Foundation and others have pointed to studies showing compiled languages like Rust and C using an order of magnitude less energy than interpreted languages for equivalent workloads. Just as F1 teams discovered that fuel efficiency constraints forced them to build fundamentally better engines, switching to Rust forces you to think about memory layout, allocation patterns, and data flow in ways that Python’s garbage collector lets you ignore – until the bill arrives.

And here is the irony that would make an F1 sustainability officer wince: Earth observation processing is meant to monitor the planet’s health. Burning excess energy to do it is like running your emissions-monitoring car on leaded fuel. F1 recognised that the sport’s 20-car grid is only 1% of its total carbon footprint, but pursued fuel efficiency anyway because the technology trickles down. The same logic applies to EO processing pipelines. The individual savings per scene are modest, but at continental archive scale they compound – just like how F1’s hybrid innovations now power road cars from Ferrari’s SF90 to the electric components in every modern turbo engine.

Static memory profiling makes this tangible. In Rust, I can tell you at compile time that a Sentinel-2 full-scene atmospheric correction will peak at approximately N gigabytes of memory, because the allocations are deterministic. In Python, you find out at runtime – usually when the OOM killer visits your pod. F1 teams know their fuel load to the gram before the formation lap. Rust gives you the same certainty for compute.

Kubernetes 1.35 and Vertical Pod Autoscaling

This deterministic memory behaviour dovetails nicely with Kubernetes 1.35’s improvements to Vertical Pod Autoscaler (VPA). VPA watches your pod’s actual resource usage and adjusts CPU and memory requests/limits accordingly. When your workload has predictable resource usage – as Rust workloads tend to – VPA converges quickly to the right allocation instead of oscillating between OOM kills and wasted headroom.

For a processing pipeline that ingests satellite scenes of varying sizes (a Landsat scene is 62 million pixels across 7 bands; a PACE scene is 2.2 million pixels across 291 bands), VPA can right-size pods per sensor type. Rust’s static memory profile means the VPA recommendations stabilise fast, which means tighter bin-packing, which means more scenes processed per node, which means lower cost per scene.

Compare this to Python pods where memory usage is non-deterministic, garbage collection spikes are unpredictable, and the VPA has to overprovision to avoid OOM. The 2 GiB/core cloud ratio that Rob’s acolite-mp was carefully designed around becomes less of a constraint when your language does not waste half of it on interpreter overhead.

Out-of-Band Development: Preventing Merge Conflicts with Upstream

One design decision I am particularly happy with is keeping the Rust port on a separate feature branch (feature/rust-port) and treating it as out-of-band from the Python codebase. ACOLITE upstream is actively maintained by Quinten Vanhellemont at RBINS, with regular additions of new sensors, algorithm refinements, and bug fixes. A traditional “rewrite in Rust” approach would create an immediate fork that diverges with every upstream commit.

Instead, the Rust code lives in src/, benches/, and tests/ directories that do not exist in upstream Python ACOLITE. The Python code in acolite/ stays untouched. The regression tests are the synchronisation mechanism – they import both the Python ACOLITE modules and the compiled Rust binary, run the same scene through both, and compare outputs.

When upstream adds a new sensor or changes a gas transmittance coefficient, the regression tests fail in the Rust port. That failure is the trigger: it goes into the agent harness as a --task, Kiro investigates the numerical difference, Copilot cross-references the upstream commit, and the fix lands in Rust without touching a single Python file. No merge conflicts. No rebasing nightmares. Just tests that enforce parity.

This is how you keep an acceleration layer in sync with a moving target – you do not try to merge them. You test them against each other.

What Is Next: The Gap to Full Sensor Parity

The roadmap has the current state at 48 Rust tests, 141 Python regression tests, and three sensors fully validated (Landsat 8/9, Sentinel-2 A/B, PACE OCI). The architecture – loader, AC, writer – is clean and extensible. But three sensors out of 30+ is a qualifying lap, not a race win. Here is what closing the gap to full ACOLITE parity actually looks like.

Sensor Coverage: 3 down, 30+ to go

Python ACOLITE supports a sprawling constellation of sensors. The Rust port has ticked off the three highest-priority ones but the remaining fleet breaks into tiers:

Tier 1 – Near-term (shared loader patterns exist):

Sensor Bands Loader Type Blocker
Sentinel-3 OLCI 21 NetCDF Sensor def exists, needs full pipeline
PRISMA 239 HDF5 Shares pattern with PACE
DESIS 235 HDF5 Shares pattern with PACE
EnMAP 224 HDF5 Shares pattern with PACE
EMIT 285 NetCDF Similar to PACE OCI

These are the low-hanging fruit. The PACE port proved out the NetCDF and hyperspectral GeoZarr writer path; PRISMA/DESIS/EnMAP share the HDF5 loader pattern. Each is a well-scoped --task for the agent harness – Kiro writes the loader and wires up the AC pipeline, Copilot validates against Python output on a reference scene.

Tier 2 – Medium-term (new loader work required):

Sensor Bands Notes
Landsat 5 TM / 7 ETM+ 7-8 Older calibration metadata formats
PlanetScope (Dove/SuperDove) 4-8 Commercial format, GeoTIFF based
WorldView-2/3 6-29 Multi-resolution, pan-sharpening
Pleiades 5-7 DIMAP format
QuickBird-2 5 Legacy but still used
VIIRS (NPP/J1/J2) 22 Swath-based HDF5, three platforms
Aqua/Terra MODIS 36 HDF4/HDF-EOS
GOCI-2 12 Korean ocean colour mission

Each of these needs a dedicated loader – different metadata formats, different calibration approaches, different file layouts. The atmospheric correction core (DSF, gas transmittance, Rayleigh, aerosol models) is shared, but getting the radiometrically calibrated top-of-atmosphere reflectance array into the pipeline is the per-sensor work.

Tier 3 – Geostationary and niche (lowest priority for aquatic applications):

GOES ABI, Himawari AHI, MTG-I FCI, SEVIRI, Sentinel-3 SLSTR, AMAZONIA-1 WFI, CHRIS, HYPERION, HICO, HyperField, HYPSO, Tanager. Some of these (HYPERION, HICO) are decommissioned but their archives are still processed. Others (Tanager at 420 bands, HYPSO at 120) are newer hyperspectral missions that would benefit most from Rust’s performance advantage.

Beyond Loaders: The Algorithm Gap

Sensor parity is not just about reading files. Python ACOLITE has several processing features the Rust port does not yet implement:

  • ROI subsetting: Limit processing to a bounding box or polygon – critical for operational workflows that do not need a full scene
  • Ancillary data retrieval: NCEP ozone, pressure, and wind speed from NASA OBPG; currently the Rust port uses default values
  • DEM-derived pressure: Copernicus DEM at 30/90m for surface pressure estimation in mountainous coastal regions
  • Glint correction: Sun glint removal for low-latitude ocean scenes
  • RAdCor adjacency correction: The physics-based adjacency effect correction developed under the STEREO program
  • TACT thermal processing: Surface temperature from Landsat thermal bands via libRadtran – this one is architecturally interesting because it requires calling an external Fortran radiative transfer code
  • Interface reflectance (rsky): Sky reflection correction at the air-water interface
  • L2W water products: Chlorophyll-a (OC algorithms), TSS (Nechad, Dogliotti), turbidity, Secchi depth – the derived products that downstream scientists actually use

The L2W gap is the most consequential. Most ACOLITE users do not care about surface reflectance per se; they want chlorophyll maps or turbidity time series. Until the Rust port can produce L2W outputs, it remains a fast atmospheric correction engine rather than a complete aquatic remote sensing toolkit.

The Realistic Path

Closing this gap is not a sprint, it is an endurance race – appropriately enough. The agent harness makes each sensor port a repeatable, testable unit of work. The pattern is established: write loader, wire to AC pipeline, run regression tests against Python, fix deltas, validate on real data. Each sensor port takes the agents a day or two of focused work plus human review.

At the current pace, Tier 1 sensors are within reach in the near term. Tier 2 will follow as the loader library matures. The algorithm features (ancillary data, glint, TACT) are orthogonal to sensor coverage and can be developed in parallel. L2W is the final milestone – when the Rust port can ingest a Sentinel-2 scene and produce a chlorophyll-a map that matches Python to within measurement uncertainty, the port will be race-ready for production.

Each of these is a --task for the agent harness. Two drivers, one constructor’s championship. The lead driver pushes into unfamiliar sensor territory, the second driver validates against Python, and the telemetry overlay catches every divergence before it compounds into a retirement.

If the intersection of Rust, Earth observation, and AI-assisted development interests you, the code is all on GitHub. Feel free to ping me with ideas, bug reports, or competing approaches – especially if you have a cleverer way to handle the N-dimensional LUT interpolation. That one was a fun 3 days of Rapid Rust Rewrite fuelled by AI Amphetamine Analogs.