A Categorical Dataset Naming System for the FCC datasets
Juraj Smieško (CERN)
IT-FCC Computing Resources Meeting
CERN
3 September 2026
join()
as the monoidal product, names as string diagrams
vertical physics pipeline
horizontal data algebra, file-level only
range(a,b) — chunk boundariesjoin(A,B,...) — monoidal product ($\otimes$)drop([c],D) — prune corrupted chunkslabel(x) — opaque registry-resolved refgen, rec, …) maps to an
append-only set of registered signatures;
the parser tells them apart structurally (argument count, literal shape),
never by resolving a value first
DetectorVersion, Tune, …)
— every parameter is typed at registration; a type can evolve
on its own without touching the signature
overlay's
two background pools) are told apart by fixed position
Overlaying background events
mc_production:key4hep_v2026>sim(IDEA_o2_v03, full,
overlay(label(bkg_pool_A), gen(10k, ZH, 240GeV, whizard, tune_A)))
Scaling statistics via monoidal joins
mc_production:key4hep_v2026>sim(IDEA_o2_v03, full, join(
gen(10k, ZH, 240GeV, whizard, tune_A, range(1, 100)),
gen(10k, ZH, 240GeV, whizard, tune_A, range(101, 200))))
Recovering from a corrupted chunk
mc_production:key4hep_v2026>drop([42], rec(IDEA_o2_v03, standard,
sim(IDEA_o2_v03, full, gen(1M, ZH, 240GeV, whizard, tune_B))))