Nine released models, one comparison set: the same 8 held-out drums
reconstructed by every model through its own identity pathway, plus each
model's signature behavior. Sources are the SAME-S codec decode of the real
sample (the fidelity ceiling every model works under); the raw original
file is included alongside, so the codec's own cost is audible too. Latest: v6,
the reconstruction model — same knobs and codes as v5, recon quality
approaching the codec floor (scoreboard in its section below). Models & code:
hf.co/lyosha/CalciumFM ·
github.
Reconstruction grid
Each model conditions on the source (mean-latent / attrs /
learned code / code+CLAP as appropriate), fresh noise, CFG 1.5, 32 Heun steps.
Listen down a column for one drum across models; expect fidelity roughly
v1.4 ≈ m1_v3 > v2 > v3.1 ≈ v3 > v5 > v4 > v4.1.
model
kick
MO_SV1_kick_aether
snare
restrange echorus snare 3
clap
8E-Clap-01
hat
WRK_10'_Minihat_Szld_1
tom
restrange tony royster jr
cymbal
Cymatics - Lofi Ride - 02
long
KSHMR War Horn 10 (C)
wild
CajintoQuiet1B
original (raw file)
source (SAME-S codec)
v1.4
m1_v3 (22M)
v2
v3
v3.1
v4
v4.1
v5
v6
v8
Signature behaviors
v1.4
The quality reference: highest fidelity, no knobs — control is only (mean latent, length) morphing + exact inversion.
m1_v3 (22M)
The 22M deployment model (50 ms/drum on 8 CPU threads); very close to v1_4 in sound.
v2
Length knob commanded −2σ vs +2σ at default guidance: the two clips are nearly identical (the style vector outvotes the knob — pooled ρ≈0.06). Same +2σ with boosted per-knob guidance: now it lengthens.
len−2 (default)
len+2 (default)
len+2 (boosted)
v3
Invert the kick's noise, then command +2σ attack spread: the edit does nothing (leaky free dims + noise carry everything). Identity is perfect but the model is uneditable this way.
invert cycle
+2σ spread (edit dead)
v3.1
The same invert-and-edit on the same kick: the spread edit lands (single hit becomes a wider burst) while the sound stays recognizably the same kick — the adversarial disentanglement at work.
invert cycle
+2σ spread (edit lands)
v4
Four free samples from perturbed library CLAP embeddings — no real sample conditioned, KAD 0.63 (conditioned band ≈0.3). Then the kick with length −2σ vs +2σ: nearly identical — the frozen CLAP outvotes the knobs. v4 samples; it does not edit.
free 1
free 2
free 3
free 4
kick len−2 (dead)
kick len+2 (dead)
v4.1
Same free-sampling recipe (CLAP + anchor knobs, KAD ≈0.97) AND the same kick length edit now responds — the scrubbed CLAP projection no longer outvotes the knobs. The one model that both samples and edits; recon is its weak point.
free 1
free 2
kick len−2 (works)
kick len+2 (works)
v5
The same kick, length −2σ / 0 / +2σ, same noise — first on v3.1 (hear it drift toward rumble as it lengthens), then on v5 (it stays the same kick, just longer/shorter). Below: v5's new F0 knob at −1.5σ / 0 / +1.5σ — the kick retunes. v5's trade: slightly less polished overall texture (KAD 0.52 vs 0.32).
v3.1 on the same kick / same noise (the before):
kick len−2
kick len 0
kick len+2
v5 (the after):
kick len−2
kick len 0
kick len+2
F0 −1.5σ
F0 0
F0 +1.5σ
v6
v6 (released) = v5 + 40k decoder-only polish steps with control pinned by self-distillation from the v5 teacher; encoder frozen, so v6 codes are drop-in with v5's. Grid: 12 sources x the full editing repertoire. Sources marked * were picked from the top of the measured v6-v5 recon-delta distribution — the median per-sample gain is +0.01 cos (subtle on short, easy drums; both models are near the ceiling there), the audible gains concentrate in hard samples like these. † recon at CFG 2.0 + z0x0.7, SAME noise for v5 and v6. Edits run at CFG 2.0 with full-temperature z0 (shrunken noise kills structural knobs — if a knob feels dead in your editor, check the z0 temperature first). PITCH NOTE: a 45 Hz kick sits at the dataset's pitch FLOOR (encoded f0 anchor z=-1.4), so f0 -1.5sigma barely moves it (45->43 Hz) while +1.5sigma retunes it clearly (45->100 Hz) — pitch-down works on mid/high-pitched sources (snare, tom, hat), pitch-up works everywhere. Length edits are RELATIVE, +-0.75 sigma around the encoded length of the source (CFG 1.5): moderate, natural envelope changes (bigger jumps push into synth-like tails). The last column is the recommended identity-preserving structural edit: invert the source, blend 50/50 with fresh noise, regenerate with the edited code.
recon-pathway metric
v5
v6
latent cos, default (CFG 1.5)
0.789
0.823
latent cos, tuned (CFG 2.0, z0×0.7)
0.840
0.859
recon-KAD, default (real-data floor 0.097)
0.523
0.322
recon-KAD, tuned
0.281
0.186
kick attack drift over the length sweep
−0.22 oct
−0.25 oct
attack drift, 16-style median
0.91
0.34
len / f0 knob pooled Spearman
0.66 / 0.74
0.56 / 0.65
invert-edit identity
0.97
0.97
source
original
SAME-S codec
v5 recon†
v6 recon†
len −0.75σ
len +0.75σ
f0 −1.5σ
f0 +1.5σ
energy +1.5σ
SDEdit s=0.5
invert-edit len+1.5σ (α=0.5)
kick
snare
clap
hat
tom
cymbal
long
wild
cdplayer *
kick2 *
cymbal2 *
longfx *
v6-MoE — per-cluster LoRA adapters (prototype)
8 rank-16 adapters over frozen v6, one per CLAP timbre cluster, control pinned by self-distillation (adapters on HF). The win is texture realism: routed recon-KAD 0.186 → 0.135 at the tuned point (real-data floor 0.097); median latent-cos gain is small (+0.001–0.003), so these five sources are picked from the TOP of the measured per-sample delta (+0.02–0.05) — where the difference is audible. Same noise for both recons; editing knobs verified intact (len ρ 1.0, invert-edit 0.97).
source
original
SAME-S codec
v6 recon (default)
v6 recon (tuned)
v6-MoE recon (tuned)
KSHMR Weird Kick 03 (G)
Abroxis - Meta Kick 17
Agogos2A
Kai Whiston - Sample Pac
KSHMR Short Fill 12 - 12
v8
v8 makes the INVERSION pathway editable: the VST encodes a sample, inverts it once (identity 0.994 at CFG 1.5 — hear the recon column), then edits knobs directly at that noise. Before v8, a length edit at full identity destroyed the sound (v6 column: kick attack drift -3.7 oct, 8/24 styles dead). v8 was fine-tuned on inversion couplings (pair-reflow): same edit now lands cleanly (+0.1 oct, 0/24 dead, aggregate drift 0.26 — better than any prior model at random noise). Last column: pitch +1.5σ at full identity. All edits are RELATIVE, in z-score units of the library distribution: f0 sigma = 1.64 octaves, so +-1.5 sigma retunes ~+-2.5 octaves from the source pitch (clamped to the +-2.5 sigma training range); scatter = transient count (rolls/claps vs one hit), spread = attack smear.
v9.1 — real-pair pitch (experimental)
v9.1 (experimental, not released): the pitch knob retrained WITHOUT any DSP-processed audio — 81,177 natural pairs mined from sample-pack note families (same instrument, different note). Realized pitch span 2.26 -> 3.63 octaves vs v6 (rho 1.0). Rows: f0 -1.5σ / encoded / +1.5σ, same noise, v6 vs v9.1. Length knob currently needs curated pairs (dead on most styles) — next calibration step.
source
v6 f0 −1.5σ
v6 encoded
v6 +1.5σ
v9.1 f0 −1.5σ
v9.1 encoded
v9.1 +1.5σ
kick
snare
tom
source
original
SAME-S codec
v8 recon (inverted, CFG 1.5)
len +1.5σ v6 (before)
len +1.5σ v8
len −0.75σ v8
f0 +1.5σ v8
f0 −1.5σ v8
scatter +1.5σ v8
scatter −1.5σ v8
spread +1.5σ v8
kick
snare
tom
cdplayer
longfx
All metrics, protocols and the full research journal: see the repo. Built from the released checkpoints.