# Drum Timing Tests: 56% Is Continuous—Test Timing-Head Updates

Evelyn Porter · September 27, 2026

> The claimed 56% drum-timing result lacks evidence. Learn why 56% makes the timing head the first audit target and how to test output-head updates.

| Takeaway | Detail |
| --- | --- |
| 56% is underdefined. | No fetched source supplies a denominator, baseline, comparison condition, or calculation for 56%, so it cannot carry the headline’s causal weight. |
| 56% is not documented as a drum result. | The corpus contains no drum-performance protocol, recordings, MIDI, sheet music, onset deviation, or timing-variance evidence tied to 56%. |
| 56% makes the timing head the first audit target. | Test output-head updates against the unchanged path; a miss that follows the output change points toward emitted timing rather than automatically proving that the groove representation was erased. |
| 56% does not prove precision-specific damage. | Only a matched FP16 ablation reproducing the same failure under fixed task, data, evaluation, and update conditions would justify the stronger claim; the accessible retraining repository never mentions 56%. |

56% is the only concrete performance figure in the supplied record, yet no fetched source defines its denominator, baseline, comparison condition, or calculation. That absence matters: the number can frame a timing question, but it cannot independently establish what failed. Continuous should therefore be treated as a hypothesis about update behavior, not as a documented result.

The first test should target the output timing head. A missed boundary can arise because the model’s emitted hit placement drifts even when the underlying representation remains intact. The supplied corpus contains no drum recordings, MIDI, sheet music, onset measurements, or rhythm protocol, so it cannot distinguish an output-head problem from a detector-threshold problem or a mismatch between a reported score and actual groove.

A matched FP16 ablation is the needed stronger comparison: hold the task, data split, evaluation procedure, and update conditions constant, then compare whether the same failure appears. Only replication of that kind would justify tying 56% to precision-specific weight behavior rather than the timing head. Until then, the defensible conclusion is narrower: audit timing-head updates, report the missing denominator and baseline, and seek evidence that can separate model output from groove-level performance.

![Drum Timing Tests](https://static.mm-ais.com/article-images-ai/drum-timing-tests-56-is-continuous-test-ai-7daa088a.jpg)

## 56% Is a Continuous Ratio, Not a Four-Bit Grid

“56%” is a normalized interval share, not a tempo multiplier. I define 56% as r = t1 / (t1 + t2), where t1 and t2 are adjacent eighth-note inter-onset intervals. At r = 0.56, the first interval occupies that share of the pair; tempo has not changed by the same proportion. Dividing by the pair’s total duration separates local long–short feel from global speed: a passage can accelerate while preserving its swing ratio. The article headline supplies the target but no denominator or measurement protocol, so this is an explicit evaluation definition, not a claim of external validation.

The second trap is categorical. I separate rhythm quantization—snapping notes to a musical grid—from low-bit neural quantization, which maps selected weights or activations to a 16-value four-bit code. This guide tests the latter. Grid snapping can erase the very local asymmetry being measured, whereas neural quantization approximates values inside the model. Confusing the two would let a labeling artifact masquerade as a four-bit failure. The evaluation must therefore retain continuous timing targets and change only the intended neural-quantization path.

I pass the previous and next inter-onset intervals, tempo, style, beat phase, and subdivision through a Transformer timing head. It predicts onset class together with continuous r. The surrounding intervals expose local timing trajectory; tempo constrains scale; style, phase, and subdivision condition the expected feel. A straight-grid prior supplies context, but it is not a hard ceiling. That distinction lets the head favor measured local asymmetry instead of snapping every prediction back to equal eighth notes.

I train this head with Huber loss for r and cross-entropy for onset and beat class. Huber loss prevents a small set of unusually distorted intervals from dominating ratio learning, while cross-entropy preserves discrete event and beat decisions. I oversample pairs near the target ratio and weight them above generic straight or triplet examples, making borderline cases central rather than incidental. The sampling proportion and loss weight should be disclosed because their effects vary with the performer-held-out set.

I score median absolute swing-ratio error separately from onset macro-F1 at a fixed millisecond tolerance. The ratio metric measures timing placement; the detection metric measures whether hits remain present. Both are computed on the same held-out performances with matched decoding conditions. A reduction in ratio error cannot conceal lost detections: if the timing head learns the feel by suppressing events, it has not produced a defensible adapter.

If the four-bit model misses the target, I retrain only the timing head with a low-rank residual, leaving the backbone untouched. The adapter competes with FP16 under two conjunctive gates. “Within 1 point” means an absolute difference in macro-F1 points, not a relative percentage. Passing both gates earns deployment; missing either one makes FP16 the defensible fallback.

| Timing-head update | Median absolute swing-ratio error | Onset macro-F1 at ±10 ms versus FP16 | Deployment decision |
| --- | --- | --- | --- |
| Both gates met | ≤2 percentage points | Within 1 point of FP16 | Deploy the low-rank adapter over the four-bit backbone. |
| Either gate missed | >2 percentage points or outside the acceptable result | Outside the 1-point band or unusable detections | Use FP16 rather than retraining the full network. |

![56% Is a Continuous Ratio, Not a Four-Bit Grid — Drum Timing Tests](https://static.mm-ais.com/article-images-ai/drum-timing-tests-56-is-continuous-test-ai-9c6c44cc.jpg)

## Large-Model Feasibility and Timing-Head Updates

As of 2026, the defensible reading is narrow: these papers justify testing a localized four-bit repair, but they do not validate a drum-timing deployment. The supplied records do not connect their model-scale and quantization results to performer-held-out drum pairs. I use them as feasibility evidence—constrained adaptation, localized learning, and lower-cost inference—not as substitutes for the timing gates.

I read Dettmers et al.’s QLoRA paper (NeurIPS 2023) as feasibility evidence, not transfer evidence. Its NormalFloat and double-quantization setup shows that a quantized large-model backbone can be fine-tuned within constrained memory while preserving quality on the evaluated benchmarks. The transferable mechanism is that low-bit storage can leave room for a task-relevant residual; those benchmarks do not establish performer-held-out swing or onset accuracy.

Hu et al.’s LoRA paper (ICLR 2022) supports the update-placement decision. Its language-model evidence favors changing a small low-rank subset instead of optimizing every weight. For this problem, I would freeze the backbone and attach the residual to the timing head, preserving the learned representation while concentrating capacity on inter-onset structure. That is architectural support for the repair, not evidence that it will generalize.

Frantar et al.’s GPTQ paper (ICLR 2023) supplies the systems case. Its large-model result makes memory reduction and A100 speedup relevant to the cost of experimentation, but neither metric measures timing fidelity. A faster quantized model can still miss the musical target. GPTQ therefore earns a benchmark slot; it does not earn deployment approval.

Henk Honing’s *The Perception of Swing in Music* supplies the semantic check. Honing distinguishes straight subdivision from triplet swing; under his interval-share convention, the already stated target is neither label. It must be stored as its own measured ratio, or the model could optimize familiar triplet timing while evaluation appeared to reward the requested condition. That labeling error would be invisible to memory and speed metrics.

For the repair itself, LoRA’s localization principle wins: QLoRA establishes feasibility, GPTQ motivates benchmark cost, and Honing keeps the target definition honest. When the four-bit model misses, retrain only the timing head with a low-rank residual—not the full network. The article’s rule supplies acceptance: median absolute swing-ratio error must be no greater than two percentage points, and ±10 ms onset macro-F1 must stay within one point of FP16. If either condition fails on performer-held-out pairs, retain FP16.

| Evidence | Reported result | Action in this decision | Verdict |
| --- | --- | --- | --- |
| QLoRA | According to Dettmers et al. (NeurIPS 2023): 65B parameters; one 48 GB GPU; 4-bit NormalFloat with double quantization; quality preserved on evaluated benchmarks. | Test whether a timing-specific residual can be fitted. | Wins the feasibility case; timing quality remains unproved. |
| LoRA | According to Hu et al. (ICLR 2022): fewer trainable parameters than full fine-tuning in language-model experiments. | Freeze the backbone; fit a low-rank timing-head residual. | Wins the update-localization case. |
| GPTQ | According to Frantar et al. (ICLR 2023): about four GPU hours; 80 GB to less than 5 GB; 3.25× A100 inference speedup at 4-bit. | Use the result to justify a low-bit inference benchmark. | Wins the efficiency case, not the accuracy case. |
| Swing definition | According to Henk Honing: 1:1 straight subdivision; 2:1 triplet swing; 66.7% in the first interval. Applying that convention to the already stated target gives 1.27:1. | Store the measured target separately from triplet shorthand. | Wins the target-definition case. |

![Large-Model Feasibility and Timing-Head Updates — Drum Timing Tests](https://static.mm-ais.com/article-images-pixabay/drum-timing-tests-56-is-continuous-test-54c90ce6.jpg)

## Timing-Head Adaptation Wins the 56% Retraining

**Automatic retraining is not an optimization strategy; it is only a trigger.** The vishakha2121/Enterprise-AI-Continuous-Learning-Platform repository describes automatic model retraining, but the supplied corpus contains no primary test report establishing timing accuracy, memory use, or a causal benefit from retraining a quantized network. Accordingly, “Not measured” below is an evidence boundary, not a disguised estimate. “Reference” identifies FP16 as the comparison anchor without inventing its score.

| Method | Trainable scope | 56% ratio error | Onset macro-F1 | Peak memory | Decision |
| --- | --- | --- | --- | --- | --- |
| Frozen four-bit | None | Not measured in supplied record | Not measured in supplied record | Not measured in supplied record | Keep unchanged if it passes |
| Four-bit + low-rank timing-head adapter | Timing head only | Not measured in supplied record | Not measured in supplied record | Not measured in supplied record | Explicit winner when the frozen model fails |
| Frozen FP16 | None | Reference value not supplied | Reference value not supplied | Reference value not supplied | Quality baseline and fallback |
| Four-bit + full-network retrain | All weights | Not measured in supplied record | Not measured in supplied record | Not measured in supplied record | Reject by default |

**Winner: retrain the timing head only; do not retrain the entire quantized network.** The adapter is a low-rank residual that leaves the backbone and other network components frozen. This constrains the hypothesis being tested: whether localized timing supervision can repair the observed ratio error without sacrificing onset detection or inheriting full-network memory costs.

I run every row on identical performer-held-out pairs, audio context, decoder settings, and seed set so quantization effects are not confounded with data or decoding changes. Pair membership, context boundaries, preprocessing, decoder parameters, and seeds must be locked before evaluation. Otherwise, an apparent quantization effect could actually be a split, context, or search-parameter effect.

I rank candidates lexicographically: lower median absolute swing-ratio error first, higher onset macro-F1 second, and lower peak inference memory third. Quality is therefore not collapsed into an arbitrary weighted score. After the frozen four-bit model misses the target, the adapter is deployable only if it reaches no more than two percentage points of median absolute ratio error while remaining within one point of FP16 on ±10 ms onset macro-F1. Failure of either gate sends the decision directly to FP16.

I report trainable-parameter count, gradient memory, optimizer-state memory, and peak inference memory separately. Trainable parameters identify what changed; gradient memory captures allocation during backpropagation; optimizer state captures persistent update storage; and peak inference memory measures deployment requirements. A small adapter share therefore does not automatically establish a small runtime footprint. The supplied record provides no defensible values for these quantities, so none are fabricated here.

I retain full-network retraining only as a falsification control. If timing-head adaptation cannot improve the target error after labels, performer splits, and interval alignment are checked, the diagnosis shifts toward representation or data rather than insufficient optimization epochs. That failure does not authorize default whole-network retraining. When the localized adapter still cannot satisfy both deployment gates, FP16 remains the defensible fallback.

![Timing-Head Adaptation Wins the 56% Retraining — Drum Timing Tests](https://static.mm-ais.com/article-images-pixabay/drum-timing-tests-56-is-continuous-test-a76dacdb.jpg)

## What the Data Doesn't Tell You

A low-rank timing-head update is not evidence of swing accuracy merely because it is cheaper or preserves a broad metric. Model size, language-model loss, and GPU-memory use describe capacity or resource behavior; they do not measure drum-onset error. The supplied material contains no drum audio, MIDI, sheet music, quantized model version, numerical format, hardware target, or before-and-after drum metric. I therefore require a 56%-specific audio/MIDI evaluation on performer-held-out pairs before calling a model swing-accurate.

A pooled target mean can also be an averaging artifact. Opposite errors across drummers, tempos, and styles can cancel, leaving a central value that describes no actual performer well. Before declaring the target learned, I require per-performer ratio distributions and subgroup results, including medians and dispersion rather than only a pooled mean. A favorable aggregate paired with directionally opposed subgroups is unresolved evidence, not successful acquisition.

No supplied snippet reports BPM, milliseconds, onset deviation, timing variance, sample size, or a drum-specific test procedure, so no numerical sensitivity example is supportable from the ledger. Annotation tolerance, onset quantization, or human relabeling can still resemble a model change. Adapter and baseline outputs should therefore be rescored against identical pair boundaries and the same annotation policy.

The diagnostic question is not merely whether aggregate performance improves, but whether the target subset moves with it. When target examples are sparse, easier bins can carry a headline score while the behavior that motivated the update remains weak. Subset-level error and recall must therefore remain visible instead of being dissolved into one pooled number.

These limits make the update rule conditional. If the frozen four-bit model misses the target, I update only the timing head with a low-rank residual—not the whole network. The adapter is eligible only after it clears both stated numerical gates on performer-held-out data; a missing, unstable, or inconclusive result is not evidence that those gates were met.

| Diagnostic condition | Defensible interpretation | Required action |
| --- | --- | --- |
| Target-band pairs constitute a small share of examples, and overall macro-F1 rises while target-band recall falls. | The aggregate improvement does not establish target learning. | Report target-subset recall and ratio error separately; do not substitute overall macro-F1. |
| A paired bootstrap’s confidence interval crosses zero, or the winner changes across seeds or performer splits. | The comparison is inconclusive. | Do not attribute the difference to retraining; report seed and split stability separately. |
| Blind drummer preference favors the frozen model despite better ratio error. | Numerical timing and perceived groove validity are separate endpoints. | Treat groove validity as unresolved and preserve both results. |
| The low-rank timing-head adapter clears the stated ratio-error and FP16 onset macro-F1 gates on performer-held-out data. | The adapter is numerically eligible under the canonical rule. | Deploy only the low-rank timing-head adapter; otherwise use FP16. |

![What the Data Doesn&#039;t Tell You — Drum Timing Tests](https://static.mm-ais.com/article-images-pixabay/drum-timing-tests-56-is-continuous-test-346ac4b6.jpg)

## Claimed Drum Performances: A 56% Case Without a Reproducible Protocol

Any reproducible denominator would be the post-filter, performer-held-out pair set—not the corpus total. The article cites Duan et al.’s Groove MIDI Dataset from *A Neural Network for Groove Modeling* as the intended source, but the supplied record does not verify its performance, duration, or measure counts. Any such totals would be provenance figures; none is the number of eligible target pairs.

I extract adjacent onsets only when both events fall within the same metrical subdivision and measure, calculate *r* for every candidate pair, and retain only 0.54 ≤ *r* ≤ 0.58. I then make a 70/15/15 split by performer identity, keeping sessions intact when session information is available, so no performer crosses partitions. I publish the post-filter pair count with the exclusion ledger and split manifest rather than treating the corpus total as the target sample size. That distinction prevents dataset size from masquerading as evaluation power.

The supplied record does not verify a tempo, beat duration, interval pair, or straight-grid displacement, so this article does not present tempo arithmetic as a measured case.

As a clearly labeled arithmetic sanity check—not a published model result—a finite low-bit grid can place a continuous target on a nearby discrete level. Such an exercise exposes only a possible mismatch an adapter would need to correct; it does not demonstrate that an adapter learned the correction.

I then run frozen four-bit, timing-head-adapted four-bit, and frozen FP16 models on exactly the same held-out pairs. For each run, I record median absolute swing-ratio error, onset macro-F1 at a ±10 ms tolerance, peak memory under identical measurement conditions, and per-performer spread. Per-performer distributions accompany pooled results so a satisfactory aggregate cannot conceal a consistently weak subgroup. Measured outcomes belong in a results ledger; the toy quantization arithmetic remains separately labeled.

The deployment gate is conjunctive. If frozen four-bit misses the target, I update only the timing head with a low-rank residual while keeping the rest of the network frozen. The adapter is deployable only if it passes both gates below. Failing either condition sends the case back to frozen FP16; full-network retraining is not the prescribed response.

| Adapted four-bit result | Median absolute ratio error | Onset macro-F1 gap versus FP16 | Decision |
| --- | --- | --- | --- |
| Both gates pass | ≤2 percentage points | Within 1 point | Deploy the low-rank timing-head adapter |
| Ratio gate fails | >2 percentage points | Measured, but cannot rescue the failed gate | Use frozen FP16 |
| Only onset gate fails | ≤2 percentage points | >1 point below FP16 | Use frozen FP16 |

![Claimed Drum Performances: A 56% Case Without a Reproducible Protocol — Drum Timing Tests](https://static.mm-ais.com/article-images-pixabay/drum-timing-tests-56-is-continuous-test-5c066426.jpg)

## Five Rules to Pass, Adapt, Reject, or Fall Back at

Retraining is not the default repair; it is a conditional intervention with a rollback path. The article metadata identifies 56% as the target, but the supplied record documents no completed quantized-model timing experiment. I therefore treat the following matrix as a predeclared deployment policy, not as evidence that an adapter has already recovered the target.

| Observed state | Required action | Acceptance test | Decision |
| --- | --- | --- | --- |
| The frozen four-bit model has held-out median absolute swing-ratio error of no more than 2.0 percentage points, while its ±10 ms onset macro-F1 is within 1.0 point of FP16. | Leave every base-network weight unchanged. Do not retrain merely to chase a marginal improvement on an already passing model. | Both quality gates pass on the performer-held-out pairs. | Keep the frozen four-bit model. |
| Held-out median absolute swing-ratio error exceeds 2.0 percentage points, but ±10 ms onset macro-F1 remains within 1.0 point of FP16. | Freeze the base network and train a low-rank residual adapter on the timing head only. | Accept the adapter only if median absolute swing-ratio error reaches no more than 2.0 percentage points and onset macro-F1 remains within 1.0 point of FP16. | Deploy only after both gates pass. |
| The adapter lowers swing-ratio error, but its ±10 ms onset macro-F1 falls more than 1.0 point below FP16. | Reject it. Inspect target-band labels and class weighting, then revise the timing loss; do not retune the entire network. | The revised candidate must clear both original quality gates. | Use FP16 while the adapter remains unapproved. |
| The adapter’s result does not reproduce across three random seeds and performer-held-out splits. | Classify the timing result as inconclusive rather than selecting the most favorable run. | The same gate outcome must recur across the required seeds and splits. | Keep FP16 as the deployment choice until the timing result stabilizes. |
| FP16 also fails either quality gate. | Collect or relabel more target-band examples, or change the timing architecture. Do not retrain the same quantized network from scratch. | Any replacement must still satisfy both predeclared gates under the held-out protocol. | Pause deployment approval rather than confounding the diagnosis with a full-network rewrite. |

Before training, I lock the performer-held-out pairs, FP16 reference, and pass/fail table so that neither the split nor the acceptance criteria can move after seeing the adapter. This rejects the status-quo myth that more retraining is automatically more rigorous: changing only the timing head preserves the model’s identity and makes a failed gate diagnostically useful.

## What to do next

| Step | Action | Why it matters |
| --- | --- | --- |
| 1 | Define 56% as the first adjacent eighth-note inter-onset interval divided by the sum of both intervals, and publish its denominator, baseline, comparison condition, and calculation. | The supplied record does not define 56%, so it cannot independently establish what failed or whether tempo changed. |
| 2 | Create a fixed drum evaluation set with recordings, aligned MIDI or verified onset timestamps, then run the four Frequently Asked Questions Does a 56% swing ratio mean the music is 56% faster? No; 56% is the normalized share r = t1 / (t1 + t2) occupied by the first of two adjacent eighth-note inter-onset intervals, not a tempo multiplier. What should be tested first if the four-bit model misses the target? Test timing-head updates against the unchanged path, because emitted hit placement can drift even when the underlying groove representation remains intact. What should be retrained when the four-bit timing model misses the target? Retrain only the timing head with a low-rank residual while leaving the backbone frozen. What two gates must the low-rank adapter pass for deployment? Median absolute swing-ratio error must be no greater than two percentage points, and onset macro-F1 at ±10 ms must be within one percentage point of FP16. What is the fallback if either deployment gate is missed? Use FP16 rather than retraining the full network. What evidence would be needed to attribute the failure specifically to four-bit precision? A matched FP16 ablation must reproduce the same failure under fixed task, data, evaluation, and update conditions. Quick answers What does the article state about the 56% figure in relation to drum timing tests? | 56% is underdefined and cannot carry the headline’s causal weight because no fetched source supplies a denominator, baseline, comparison condition, or calculation for it. |
| What is the first test recommended in the article for addressing timing issues? | The first test should target the output timing head. |  |
| How does the article define 56% in terms of timing intervals? | I define 56% as r = t1 / (t1 + t2), where t1 and t2 are adjacent eighth-note inter-onset intervals. |  |
| What condition must be met to justify tying 56% to precision-specific weight behavior rather than the timing head? | Only a matched FP16 ablation reproducing the same failure under fixed task, data, evaluation, and update conditions would justify the stronger claim. |  |
| What does the article say about the relationship between timing-head updates and emitted hit placement? | A missed boundary can arise because the model’s emitted hit placement drifts even when the underlying representation remains intact. |  |

Also worth reading: **Lo-fi drum patterns: 90 Beats Per Minute Ableton Live Freeze Wins vs Bounce**: [Lo-fi drum patterns: 90 Beats](https://getrhythmm.com/blog/lo-fi-drum-patterns-90-beats-per-minute-ableton-live-freeze-wins-vs-bounce.php) · **AI rhythm generation is rewriting how modern tracks get made**: [AI rhythm generation is rewriting](https://getrhythmm.com/blog/ai_rhythm_generation_is_rewriting_how_modern_tracks_get_made.php) · **AI Beat Making for Short-Form Video: Rhythm, DAWs, and Visual Downbeats**: [AI Beat Making for Short-Form](https://getrhythmm.com/blog/ai-beat-making-for-short-form-video-rhythm-daws-and-visual-downbeats.php)

### Related reading

- [Ableton AI Drum Generation: 10-Track Test Cuts Time 40% vs Manual](https://getrhythmm.com/blog/ableton-ai-drum-generation-10-track-test-cuts-time-40-vs-manual.php)
- [2026 A/B Test: AI Drum Loops vs DAW Patterns for Podcast Intros](https://getrhythmm.com/blog/2026-ab-test-ai-drum-loops-vs-daw-patterns-for-podcast-intros.php)
- [Lo Fi Hip Hop Kicks: 3ms Synth Wins 4-1 vs Stack at -9 Loudness Units (LUFS)](https://getrhythmm.com/blog/lo-fi-hip-hop-kicks-3ms-synth-wins-4-1-vs-stack-at-9-loudness-units-lufs.php)
- [Lo Fi Drum Swing: Ableton Live Groove Pool 57% Hold vs Fail](https://getrhythmm.com/blog/lo-fi-drum-swing-ableton-live-groove-pool-57-hold-vs-fail.php)
- [How to use AI rhythm for Twitch stream transitions and alerts](https://getrhythmm.com/blog/how-to-use-ai-rhythm-for-twitch-stream-transitions-and-alerts.php)
- [Lo Fi Hip Hop Drums 2026: 68% Prefer Artificial Intelligence Swing vs Drum Machine](https://getrhythmm.com/blog/lo-fi-hip-hop-drums-2026-68-prefer-artificial-intelligence-swing-vs-drum-machine.php)

### Latest

- [Lo Fi Hip Hop Kicks: 3ms Synth Wins 4-1 vs Stack at -9 Loudness Units (LUFS)](https://getrhythmm.com/blog/lo-fi-hip-hop-kicks-3ms-synth-wins-4-1-vs-stack-at-9-loudness-units-lufs.php)
- [Lo Fi Drum Swing: Ableton Live Groove Pool 57% Hold vs Fail](https://getrhythmm.com/blog/lo-fi-drum-swing-ableton-live-groove-pool-57-hold-vs-fail.php)
- [How to use AI rhythm for Twitch stream transitions and alerts](https://getrhythmm.com/blog/how-to-use-ai-rhythm-for-twitch-stream-transitions-and-alerts.php)

Canonical: https://getrhythmm.com/blog/drum-timing-tests-56-is-continuoustest-timing-head-updates.php
Markdown: https://getrhythmm.com/blog/drum-timing-tests-56-is-continuoustest-timing-head-updates.php/index.md
