MeterStore

The CLI

meterstore init, check, create, status, archive, maintain, query, completeness and serve — the same library, without a program to write.

cargo install meterstore --features cli

A thin front end over the same public API a service uses: every subcommand is one or two library calls, and nothing here is unreachable from Rust.

init checkWrite a starter configuration; validate it, connecting to nothing
createBoth tiers, every declared table. Idempotent
statusBoundary, lag, write runway, health. Non-zero exit when unhealthy
archive maintainOne-shot for cron; foreground loop for a sidecar
query explainSQL across both tiers, with the boundary it ran against
completenessWhich channels are short, and which delivered nothing
snapshotsWhat a settlement rerun can pin to
erasuresThe erasure audit trail, deployment-wide
serveFlight SQL and the Iceberg REST façade
purgeDestroy a table. No recovery path

Every command takes -c/--config (or METERSTORE_CONFIG) and defaults to meterstore.toml in the working directory.

Getting to a working store

meterstore init            # a commented starter configuration
$EDITOR meterstore.toml
meterstore check           # full validation — no database needed
meterstore create          # both tiers, every declared table
meterstore status

check connects to nothing, which is what makes it a CI step: a file whose settlement_lag is shorter than its archival_step fails there rather than stranding corrections below the watermark in production.

Keeping archival running

# The sidecar shape: a foreground loop, ctrl-c to stop.
meterstore maintain --interval 15m

# The cron shape: archive what is due, then exit.
meterstore archive --max-windows 8

Both are safe on every replica. One wins each table’s archive lease and the others report contention and stop, which is not a failure — see Operations.

maintain stops after the cycle in flight rather than at the signal. For a cycle mid-archival that is the difference between a clean stop and an orphaned partition the next run has to reclaim.

Snapshot expiry is opt-in (--expire-snapshots), because a snapshot is what makes a past settlement reproducible: how far back a settlement can be reproduced is a compliance decision.

So is the other opt-in job. --anonymise-after-years 3 runs the § 60 Abs. 6 MsbG sweep on the same schedule, destroying the linkage of every subject whose readings have passed the ceiling in every table:

meterstore maintain --anonymise-after-years 3 --anonymise-actor retention-job

Off by default and irreversible when on. The years are full calendar years after the year of collection, not now - 3 years: the statutory clock starts at the Schluss des Kalenderjahres, and the rolling spelling would erase a January value a year early. Needs a table declaring subject_column and a [privacy] section for the registry it resolves against — see Privacy and retention.

Asking what the store holds

meterstore query "SELECT meter_balancing_day(\"from\", sparte) AS day, SUM(value)
                  FROM readings
                  WHERE malo_id = '41373559241'
                  GROUP BY 1 ORDER BY 1"
+------------+------------+
| day        | sum(value) |
+------------+------------+
| 2026-07-19 | 412.750000 |
| 2026-07-20 | 408.125000 |
+------------+------------+

boundary  readings_versions: 2026-08-18T00:00:00Z
tiers     cold + hot
warning   this answer includes the hot window, so it is only valid for now —
          those intervals are still being corrected

The footer is the point. A SUM alone cannot say whether it crossed the tier boundary, and two identical queries a minute apart can read the same rows from different tiers. Every result carries the boundary it was computed against and whether it reproduces.

--historical reads the settled history alone, with no load on PostgreSQL, and the footer then says the answer does not change. --operational reads the recent window alone, with no Iceberg round trip.

explain shows the schema and the same provenance without running the statement — which tiers a query would touch is usually the thing worth knowing about a slow one.

A statement of - reads from standard input, so a long query can live in a file:

meterstore query - < settlement.sql

Asking whether a month is complete

A SUM over an incomplete month returns a smaller number and no reason, so the question has its own verb:

# The settlement period, named the way the market names one.
meterstore completeness --month 2026-06 --seen-since 30d
TABLE                MALO          OBIS           SPARTE  RES      EXPECTED    ACTUAL   MISSING   SURPLUS  FIRST GAP     NOTE
readings_versions    41373559241   1-0:1.29.0     STROM   PT15M        2880      2880         0         0  —
readings_versions    56789012345   1-0:1.29.0     STROM   PT15M        2880      2784        96         0  2026-06-14
readings_versions    99887766554   1-0:1.29.0     STROM   PT15M        2880         0      2880         0  2026-06-01    delivered nothing in the range

3 channel(s) over [2026-05-31T22:00:00Z, 2026-06-30T22:00:00Z) — 2 incomplete, 1 silent, 2976 interval(s) missing

The expectation is the DST-aware calendar’s, per balancing day: 92 intervals on the spring day, 100 on the autumn one, and for gas both on the Gastag rather than the Sunday. It covers every day of the range, so a channel that stopped mid-month is short by every day after — and a report over a month that has not finished reports the remainder as missing. Ask about periods that are over. --gaps-only drops the complete rows.

--month YYYY-MM is a Bilanzierungsmonat, not a calendar month. The range printed above starts at 22:00 UTC on 31 May because that is midnight in Berlin. --sparte GAS cuts the same span at 06:00 local instead. Rows are always counted against their own sparte; the flag decides where the range is cut, which one value cannot do for two — so a table holding both is reported once per commodity.

--from/--to take RFC 3339 instants for any other period, and --malo, --melo and --obis narrow the report to one measuring point, one meter or one channel — in the scan, so asking about one meter does not cost a scan of the portfolio. Each is parsed: a mistyped identifier fails at the call rather than narrowing to nothing, which on this report reads as “everything is fine”.

--seen-since is the finding a range cannot make about itself. A channel that delivered nothing produces no rows to aggregate, so it is absent from the report rather than reported as empty — unless a roster is drawn from an earlier window. There is no default, because what a roster means is master data this crate does not hold: too short and a meter read monthly looks decommissioned, too long and every terminated measuring point is a standing finding. The narrowing flags apply to the roster too, or every channel outside them would come back as silent.

This exits zero whatever it finds. A gap is a fact to triage; status is the check that fails, because a stranded row means query results are wrong rather than incomplete. For a monitoring check, read the JSON — it carries channels_reported, channels_incomplete, channels_silent, channels_unmeasurable and intervals_missing beside the rows:

meterstore completeness --month 2026-06 --format json \
  | jq -e '.channels_incomplete == 0 and .channels_unmeasurable == 0'

Both halves matter. A channel with no declared resolution has no expectation to compare against, so it reports complete without anything having been checked — a check on channels_incomplete alone passes a month nobody could judge.

More on what the numbers mean, and why missing is not expected - actual, in Completeness.

Checking a declaration against the data

Declaring a tenant discriminator as an attribute rather than an identity column is legal, writes succeed, and nothing raises an error — but two tenants then share a merge key, so one silently supersedes the other. It is the one schema mistake this crate cannot refuse, because nothing at write time can tell it from a column that genuinely is a fact about a reading.

So it is asked of the rows instead:

meterstore audit
TABLE                        COLUMN                       KEYS     REPEATED   WIDEST    SHARE
readings_versions            bilanzkreis                 84_112          37        2     0.0%
readings_versions            ingest_source               84_112      41_930        2    49.9%

REPEATED is how many merge keys carry more than one value of the column. A correction may legitimately restate an attribute, so a few are ordinary — the first row. Half of them is not: ingest_source is identity in everything but the declaration, and every key two sources both report is one of them.

A report, not a verdict: the two cases differ in degree, and only the deployment knows which its column is. It exits zero either way. --column and --table narrow it; an identity column and an unknown name are both refused, because a clean bill for a column nobody checked is the worst answer available.

It is a full group-by over the table, so run it deliberately rather than on a schedule — after a schema change, or when a number looks wrong.

Machine output

--format json renders one document per invocation, with the provenance beside the rows rather than under them:

meterstore status --format json | jq '.tables[] | select(.healthy | not)'
{
  "rows": [{ "day": "2026-07-19", "sum(value)": "412.750000" }],
  "row_count": 1,
  "provenance": {
    "watermarks": [{ "table": "readings_versions", "watermark": "2026-08-18T00:00:00Z" }],
    "tiers_scanned": ["cold", "hot"],
    "reproducible": false
  }
}

Decimals are rendered as strings, not JSON numbers. Settlement is money, and a double is not.

Exit codes

CodeMeaning
0Success
1Failed, and retrying will not help — a bad statement, invalid configuration, a refused delivery
75Failed transiently — the database was unreachable, a lock was not available. EX_TEMPFAIL, so a supervisor should try again

status exits non-zero when a table is unhealthy, which is what makes it usable as a monitoring check. The two ways a table stops working exit differently and the message names which:

  • Rows stranded below the watermark. Query results may be wrong. This is the alert.
  • No hot partition left ahead of the write frontier. The next insert fails outright. Check that archival is running.

Serving

meterstore serve --addr 127.0.0.1:50051 \
                 --catalog-addr 127.0.0.1:8181     # optional

Two surfaces, answering different questions.

Flight SQL carries the unified hot + cold view — the one surface an external client cannot assemble for itself, because the hot tier is not in the Iceberg catalogue.

The catalogue façade is a read-only Iceberg REST endpoint, and it is what a SQL-catalog deployment needs so Spark, Trino, DuckDB and PyIceberg can read the settled history straight from object storage — with this process in the metadata path only, never the data path, so readers scale in parallel and it is neither a bottleneck nor a single point of failure. A deployment already on a REST catalogue needs none of it: point engines at the endpoint it already has, which is why the flag is opt-in and Flight SQL’s is not.

Analytics over settled history belong on the catalogue, read directly. Routing them through Flight adds a proxy hop and serialises parallel reads through one process.

One ctrl-c stops both. A server still answering metadata while its query surface is down is worse than one that stopped, because a client cannot tell.

Unauthenticated, both of them. This crate has no business deciding what kind of authentication a deployment uses, so the library hands back a tonic service and an axum router to wrap in your own interceptor and TLS, and the CLI binds them bare. Bind them to loopback or a trusted network, never to a public one. Read-only either way: Flight refuses every mutating call and every statement that is not a query, and the façade has no write path at all — an external writer appending through it would put rows in the cold tier without MeterStore knowing, and nothing downstream could detect it.

Destroying a table

meterstore purge --table readings_versions --confirm readings_versions

The only operation in the crate that deletes stored readings — every partition, the catalogue entry and the data files in object storage. There is no recovery path, which is why the name has to be given twice.

The erasure trail

meterstore erasures --limit 50

# What an auditor actually asks for: a period, and one duty.
meterstore erasures --since 2026-07-01T00:00:00Z --until 2026-10-01T00:00:00Z \
                    --trigger retention

“We deleted it” is not evidence. Every erasure writes a row saying when, why, by whom and which duty it discharged — and deliberately not whose, since the natural identifier is the thing being destroyed. This is the report a regulator asks for.

The TRIGGER column is request for an Article 17 erasure and retention for the § 60 Abs. 6 sweep — different legal bases, asked about separately.

--since/--until are half-open, so consecutive quarters tile. A backwards period, an unknown --trigger and a non-positive --limit are refused rather than printing an empty trail, which would read as “nothing was erased”.

The registry is deployment-wide, so the trail is one list rather than one per table. A configuration whose tables declare no subject_column holds no mapping at all, and says so rather than printing an empty list that reads as “nothing has been erased”.

Two rows read differently. — (suppression only) in SUBJECT is a request that named an identifier this deployment held no mapping for — one that arrived before the delivery did: no linkage to destroy, and honoured by refusing the identifier from then on. An indented suppression lifted … line is an erasure carried out against the wrong subject and released, with who did it and why.

No meterstore erase

Reading the trail is the only thing the command line does with the subject registry, and the omission is the same shape as the one below.

An Article 17 request usually reaches an application’s own tables too — billing periods, quality assessments, substitute-value logs — and those must succeed or fail together with the mapping. A CLI invocation commits its own transaction and cannot enclose them, so the failure mode is the worst kind: a subject reported as erased whose derived rows survived. SubjectRegistry::erase_in and erase_all_in take a transaction the caller owns, which is where that erasure belongs. Privacy and retention →

The duty that comes due on its own is a different matter, and the CLI does run it: meterstore maintain --anonymise-after-years 3.

No meterstore append

A reading arrives as an MSCONS message, an SMGW push or a CSV a utility exports its own way, and each needs a mapping this crate has no opinion about — which OBIS code, which network operator issued the version, which Messlokation. That is Writing readings; check is what the CLI offers it.