RFC-0010: Decouple control/observation vector operations from datavect via a dedicated operator transform#
Field |
Value |
|---|---|
Author(s) |
Antoine Berchet |
Status |
Draft |
Created |
2026-07-29 |
Decided |
|
Reviewers |
|
Supersedes |
— (follow-up to ADR-0009, does not replace it) |
Superseded by |
— |
Summary#
CIF’s control vector is necessarily gridded today, while the observation side already
supports sparse data — an asymmetry that blocks inherently sparse/point-like control
vectors (land-process-model parameters, point sources). Both vectors are wired to
datavect (per ADR-0009), which only
recognises data the user explicitly declares in YAML: an observation operator’s internal
intermediate data has no way to be configured, and observation vector entries built from
several data points (“super-observations” — horizontal lidars, SVD/encoder-compressed
representations) can’t be represented at all. This RFC proposes splitting today’s
transform pipeline into three explicit layers — a datavect I/O transform, a merged
control/observation operator transform with explicit forward and adjoint, and a thin
fromcontrol/toobsvect interface to the raw controlvect.x / obsvect.ysim arrays —
so that sparse control vectors and explicitly-configurable intermediate transforms become
representable without depending on datavect for everything. The concrete
super-observation data structure is intentionally left to a follow-up RFC; this RFC scopes
the enabling architecture only, and is designed to be additive/opt-in so existing configs
are unaffected.
Motivation / Problem#
Control vector is necessarily gridded, observation side is already sparse. This asymmetry blocks control-vector use cases that are inherently sparse/point-like: parameter optimization for land-process models (e.g. per-site or per-PFT parameters), and point-source inversions (a handful of discrete sources, not a grid cell field).
Intermediate transform outputs can’t be configured when observation operators are complex. If transform 1 produces
data1which feeds transform 2, there is no way to specify operator-specific choices ondata1(e.g. which regridding method to apply) becausedata1isn’t declared indatavect. Per ADR-0009,obsvect/controlvect(and the auto-wiring ininit_sparse.py) only recognise what’s declared indatavect— anything not declared there has no way to be configured or referenced explicitly.Super-observations aren’t representable. A single observation-vector entry built from several underlying data points doesn’t exist as a concept today. This covers both genuinely complex observation geometries (horizontal-viewing lidars) and derived / compressed representations (SVD components, learned encoder outputs).
obsvectassumes a flat one-row-per-point sparse structure with no aggregation concept, which also blocks cluster-based analyses, satellites observed at multiple scales, geopandas-style complex observation geometries, and assimilating gridded data directly as observations.These are exactly the two “known weaknesses” ADR-0009 flagged: no element can exist in the control/observation vector without a
datavectentry, and vector membership is implicit and error-prone. This RFC is the follow-up ADR-0009 anticipated.
Context & constraints#
datavectcentralisation (ADR-0009).controlvect/standardandobsvect/standardboth declaredatavectas a hard requirement and derive their structure from it — there is no independent declaration of vector content today.Mapper / auto-wiring architecture (ADR-0002). Transforms exchange data through a shared
mapperdict keyed by(component, parameter);init_sparse.pycompares a precursor’ssparse_data/sampled/continuous_*domainflags against its successor’s and auto-inserts the needed conversion transform. This machinery currently only seesmapperentries that trace back to adatavectdeclaration.Current
fromcontrol/toobsvectresponsibilities. These system transforms today conflate two things: interfacing with the rawcontrolvect.x/obsvect.ysimarrays, and carrying out control/observation-specific operations (resolution handling, operator-specific processing). This RFC’s core move is separating those two responsibilities.Plugin interface contract (ADR-0006). Any new transform “kind” introduced here still goes through the single
Pluginbase class /plugin_typesregistry — this RFC does not touch the extensibility mechanism itself, only what happens inside thetransformplugin category.
Proposed design#
Split the pipeline into three explicit layers, replacing the current entanglement of
datavect membership with control/observation-specific operations:
datavectI/O transform — pure fetch (control-vector direction: read the prior/background data a plugin needs) or store (observation-vector direction: write data back for bookkeeping/output). No operator logic lives here.A merged control/observation operator transform, with explicit forward and adjoint — absorbs what’s split today between controlvect-side “resolution” operations and obsoperator-specific operations into one transform kind shared by both sides. Because this layer is decoupled from
datavect, its intermediate data (thedata1example above) can be explicitly configured — e.g. which regridding to use — without requiring adatavectentry for it.fromcontrol/toobsvect, narrowed to a thin interface — only the plumbing that connects the operator-transform layer tocontrolvect.x/obsvect.ysim.
This decoupling is what makes sparse control vectors and (eventually) super-observations
representable: neither is forced through the current datavect-shaped bottleneck, because
the operator transform’s internal data no longer needs a datavect entry to exist or be
configured.
Explicitly out of scope for this RFC: the concrete super-observation data structure
(how a multi-point aggregate, SVD component, or encoder output is represented inside
obsvect, and how its forward/adjoint is defined). This RFC only removes the
architectural blocker (datavect-forced 1:1 membership); the schema itself is deferred to
a follow-up RFC once this layering is in place.
Additive / opt-in: existing configs that rely on today’s implicit datavect-driven
wiring keep working unchanged. The new operator-transform layer is something a config
opts into when it needs sparse control-vector elements, explicit intermediate-transform
configuration, or (later) super-observations.
Alternatives considered#
Do nothing / keep working around it in ad hoc plugin code. Rejected: this is exactly the implicit-ambiguity failure mode ADR-0009 already flagged, and it gets worse as more complex observation operators are added.
Extend the
datavectschema instead (e.g. aninternal: trueflag letting an entry exist without being user-facing). A lighter-weight alternative that would address the intermediate-transform problem alone, without touchingfromcontrol/toobsvector enabling sparse control vectors or super-observations. Not chosen as the primary design here because it doesn’t address the control-vector asymmetry, but flagged in Risks below as worth reviewers weighing in on — it could be a smaller complementary fix or a staging step.Build super-observations as a one-off
obsvectfeature, without generalising the transform layer. Rejected: it wouldn’t solve the sparse-control-vector or intermediate-transform problems, and would likely need to be redone once those are addressed anyway.
Invariants & assumptions affected#
[x] Adjoint / tangent-linear consistency — the new merged operator transform must supply both forward and adjoint; existing gradient-check infrastructure (
check_adjtltest.py) needs to cover it explicitly.[x] Numerical reproducibility — existing reference configurations must reproduce bit-for-bit, since the change is additive/opt-in and doesn’t touch the code path configs don’t opt into.
[ ] Mass conservation / physical invariants — N/A unless the operator transform introduces new resampling; existing regrid/interpolation invariants must still hold when that logic moves into the new operator layer.
[x] Plugin interface contract — the
transformplugin category’s interface changes (narrowerfromcontrol/toobsvect, new operator-transform kind); existing transform plugins (regrid,vertical_interpolation, etc.) need to be evaluated against the new layering.[x] Config / YAML backward compatibility — by design, existing configs are unaffected; new capabilities require new YAML. No migration should be needed for current users.
[ ] On-disk / output format — N/A for this RFC; may become relevant once super-observations are designed in the follow-up RFC.
[x] Ordering dependencies / global state — the new operator layer sits between the
datavectI/O transform andfromcontrol/toobsvect;init_sparse.py’s current auto-wiring insertion order must either be preserved for the unchanged path or explicitly superseded for the opt-in path. Needs to be resolved during design, not assumed.
Impact#
Affected components / plugins / instruments:
controlvects/standard,obsvects/standard,transforms/system/fromcontrol,transforms/system/toobsvect,transforms/utils/init_sparse.py(auto-wiring),datavects/standard.Performance & memory: expected negligible if the new layer is implemented as another task in the existing Dask graph; to be confirmed during implementation.
User-facing changes: new, additive YAML syntax to declare sparse control-vector elements and to explicitly configure intermediate transforms; nothing changes for configs that don’t use the new syntax.
Documentation that must change:
docs/decisions/0009stays as-is (it accurately records the state before this RFC); new user/dev docs will be needed for the operator-transform layer once implemented, alongside the existingdatavect/controlvect/obsvecttutorials.
Validation plan#
Gradient / finite-difference adjoint checks on the new merged operator transform, reusing the existing
check_adjtltest.pyinfrastructure.Regression run on existing reference configurations to confirm bit-identical results (the opt-in guarantee above).
A small point-source or land-process-parameter benchmark exercising the new sparse control-vector path once implemented, since that’s the first concrete use case this unblocks.
Risks & open questions#
Whether the “extend
datavectschema” alternative should be pursued in addition to this restructuring, as a lighter complementary fix for the intermediate-transform problem specifically — open for reviewers.The concrete super-observation schema is deferred to a follow-up RFC; candidate use cases to keep in mind when designing it: cluster-based analyses, satellites observed at multiple scales, geopandas-style complex observation geometries, and assimilating gridded data directly as observations.
How the new operator-transform layer interacts with today’s
sparse_data/sampledflag-driven auto-wiring (ADR-0002) — replace, wrap, or coexist? Needs resolution during implementation design, not assumed here.This touches core pipeline-construction / mapper-wiring code, one of the most central and highest-blast-radius parts of the codebase — implementation should be reviewed with that in mind.
Rollout / migration#
Phased, and additive throughout — no forced migration for existing users:
Implement the
datavectI/O transform and the merged operator transform layer; existing gridded/sparse configs continue through their current (unchanged) path.Add sparse control-vector support (point sources, land-process-model parameters) as the first concrete use case exercising the new layer.
Revisit super-observations in a follow-up RFC once (1)–(2) are in place and validated.
Decision#
Outcome: Date / decider: Rationale (esp. if rejected or changed): Follow-up issues / MRs: