The canonical AUC data model¶
The canonical model is the single in-memory representation every import path converges on. It is a data-representation layer: it preserves what was supplied and makes no judgement about scientific validity or suitability for analysis. See ADR-0002.
Two layers¶
- Metadata — pydantic v2 models (frozen,
extra="forbid"): experiment identity, instrument/run, samples, and per-scan metadata. - Observations — an xarray-backed store of the numeric radial signal, built on NumPy.
AUCExperiment composes the two. It is a frozen dataclass (not a pydantic
model) so it can hold the xarray-backed Observations without forcing pydantic
to serialise an array layer it does not own.
from openauc.models import (
AUCExperiment, ExperimentMetadata, ScanMetadata, Observations, Quantity,
Unit, OpticalSystem,
)
experiment = AUCExperiment(
metadata=ExperimentMetadata(experiment_id="exp-1"),
scans=(
ScanMetadata(
scan_id="scan-1", index=0,
elapsed_time=Quantity.of(0.0, Unit.SECOND),
optical_system=OpticalSystem.ABSORBANCE,
),
),
observations=Observations.from_shared_axis(
radius=[6.0, 6.1, 6.2],
signal=[[0.10, 0.20, 0.30]],
scan_ids=["scan-1"],
signal_unit=Unit.ABSORBANCE_UNIT,
),
)
print(experiment.summary())
report = experiment.validate_structure()
assert report.is_valid
Model areas¶
| Area | Model | Required field(s) |
|---|---|---|
| Experiment identity | ExperimentMetadata |
experiment_id |
| Instrument & run | InstrumentMetadata |
none (all optional) |
| Sample | SampleMetadata |
sample_id |
| Scan | ScanMetadata |
scan_id, index, elapsed_time |
| Observations | Observations |
radius/signal arrays + scan_ids |
| Provenance | ImportProvenance |
none (all optional) |
Nominal rotor speed lives on the instrument; the actual per-scan speed and
temperature live on each ScanMetadata.
Radius axes: shared and per-scan¶
Observations are stored in one of two explicit modes (RadiusAxisMode):
- Shared — every scan shares one radius axis. The backing dataset has a 1-D
radiuscoordinate and asignalvariable with dims(scan, radius). - Per-scan — each scan carries its own radius vector. Vectors are stored in
padded 2-D arrays with dims
(scan, point)plus an authoritative boolean validity mask. Shorter scans are padded withNaNandmask=False. A value is a real observation if and only if its mask entry isTrue; padding is never presented as measured data.
Nothing is interpolated or resampled. See
missing-and-unknown-values for why the mask is
authoritative rather than relying on NaN.
obs = Observations.from_per_scan(
radii=[[6.0, 6.1, 6.2], [6.0, 6.05]], # different lengths
signals=[[0.1, 0.2, 0.3], [0.4, 0.5]],
scan_ids=["a", "b"],
signal_unit=Unit.FRINGE,
)
obs.points_per_scan() # (3, 2) — real observations per scan
obs.valid_radius_values() # excludes padding
Construction validates geometry and finiteness and raises
ObservationError on malformed input, so an invalid Observations cannot
exist.
Validation, readiness and summaries¶
Validation is tiered, and none of it raises:
report = experiment.validate_structure() # archival + structural findings
report = experiment.validate() # all four tiers
assessment = experiment.assess_readiness() # metadata presence per workflow
summary = experiment.summary_data() # structured facts
print(experiment.summary()) # the human-readable rendering
validate_structure() checks representational consistency only — identifier
uniqueness, scan/observation agreement, positive radius values, and well-defined
optical-system/signal-unit combinations. Field-level invariants (finiteness,
non-negative time and wavelength, valid masks) are enforced earlier, at
construction, and raise immediately.
Absent metadata is reported, never required: an experiment with sparse
metadata stays archivally and structurally valid, while
assess_readiness() reports separately whether a future workflow's metadata
prerequisites are present. Neither is scientific quality control, and scientific
suitability is permanently reported as NOT_ASSESSED.
See validation tiers, analysis
readiness, the source modules
openauc.models.validation, openauc.models.checks, openauc.models.readiness
and openauc.models.summary, and the API reference.
Serialisation¶
- Metadata round-trips through pydantic (
model_dump/model_validate, including JSON). Observations.to_dict()/from_dict()round-trips the numeric layer using plain Python types; per-scan padding is written asnulland the mask is written explicitly.AUCExperiment.to_dict()/from_dict()round-trips the whole experiment.
This JSON-friendly serialisation is the in-memory foundation the Phase 6 AUCX
archive will build on. Writing or reading .aucx archives is not implemented
in this phase.