# Sound workshop revision 2 — 24 September 2026

Devin rejected the first Stable Audio bank: effects sounded like dragging a rock on pavement and did not match their prompts. The earlier float-writer and latent-length corrections reduced numeric problems but did not establish creative quality. That bank is not accepted.

## What the controlled tests show

The local pinned Stable Audio 3 Small SFX setup behaves poorly on the very short source lengths we used. With the identical literal coin prompt and seed, extending generation from 1.48 to seven seconds changes a sustained low-frequency texture into separated transients. Guidance 1 alone does not rescue the short generation. This is evidence of a setup contribution, not a verdict on the model's general quality or a proven universal minimum duration.

The [publisher's Small SFX example](https://huggingface.co/stabilityai/stable-audio-3-small-sfx) uses seven seconds, eight steps and guidance 1. The revised bank follows those settings, removes the negative prompt and uses concise descriptions of concrete sounds.

| Controlled coin generation | Spectral centroid | Energy below 500 Hz | Raw peak |
| --- | ---: | ---: | ---: |
| 1.48 seconds, guidance 1 | 189 Hz | 93.1% | 0.571 |
| 1.48 seconds, guidance 3 | 256 Hz | 95.3% | 1.128 |
| 7 seconds, guidance 1 | 4,943 Hz | 1.8% | 0.784 |
| 7 seconds, guidance 3 | 3,932 Hz | 28.1% | 3.686 |

These are mono signal measurements, not listening scores. Guidance 3 tests also use the previous negative prompt, so that comparison does not isolate guidance from negative conditioning. The duration comparison at guidance 1 does isolate source duration. Seven-second clap and water controls have different timing and spectra. A verbose coin prompt still produces an overly high-frequency result; longer generation alone is not creative acceptance.

A separate SAME-S encode/decode probe reconstructed a known 440/880 Hz tone with 0.9985 correlation and 99.8% of decoded spectral energy at the target tones, for both direct and chunked decoding. This weighs against a general broken-codec explanation. It does not prove every generated latent decodes correctly.

## Revised proposals

- **26 Stable candidates:** one seven-second source per cue, eight steps, guidance 1, seed 240924, no negative prompt. Select the strongest event cluster, crop to game timing, apply fades and conservative levels. Boost uses a steady interior segment and a loop crossfade. The full source take is separately available, level-adjusted for playback; it is not an unnormalized raw master.
- **26 Astra-designed candidates:** original procedural sound designs composed in this Codex thread and synthesized locally. Tuned wood, ceramic, metal and glass resonators; pitch glides; restrained filtered air; short melodic gestures. No Astra text-to-audio endpoint was used or is available through this connection. The recipes and reproducible script identify every layer.
- **10 existing MOSS candidates:** retained for comparison. The MOSS mix uses Astra for events outside that ten-cue set.
- **Eight 11-second mixes:** original, Astra, Stable and MOSS/Astra, each dry and with existing music. These reproduce an event sequence and V3 gains, not a captured gameplay session. The shield-hit layers and crash/heart-loss layers overlap as in the game.

The direction is a warm arcade world: ceramic espresso ticks, a small metallic coin reward, buoyant rubber movement, rounded mechanical boost, glassy protection and separate physical-impact and heart-loss cues. Astra's direct synthesis provides intentional timing and distinct gestures without relying on prompt adherence. It remains a proposal, not a professional-quality claim based on measurements.

## Acceptance boundary

The assistant cannot hear audio in this connection. Finite samples, headroom, unique files and successful browser playback do not establish pleasantness, prompt fidelity, absence of unwanted speech/music, perceived loudness or seamless looping. Those remain listening judgments. The dashboard provides repetition, both loop controls and the full Stable takes to make comparison practical.

The native game's original audio remains unchanged. Choose sounds by listening before promotion; preserve Unity filenames and metadata GUIDs. No gameplay or difficulty changes belong in this audio revision.

## Reproduction and evidence

Project scripts: `Tools/Audio/refine_stable.py`, `astra_synthesis.py`, `prepare_audio_v2.py`, `stable-prompts-v2.json`, and `audition-v2.html`. Start with the reproduction section in `Tools/Audio/README.md`.

Raw generations, codec probe and controlled diagnostics are preserved on the external audio work drive under `diagnosis-v2/`, `raw/stable-v2/` and `raw/astra-designed/`, with prompts, settings, seeds and SHA-256 receipts. The project stores verification and publication receipts under `Builds/audio-lab-v2-*.json`.

The dashboard is `/audio-lab-v2/`. Its ZIP includes game-length proposals, originals, prompts, recipes and measurements; the larger full takes, diagnostic recordings and mixes are available separately online. The first workshop remains preserved at its original immutable Cloudflare deployment.
