MFPT field validation — per-file detail
# MFPT Field Validation Results
Frozen-constants cross-rig validation (Part A) and real-world field test (Part B) — Phase 7. Profile: `route` (unvalidated, unmodified for this run) — unvalidated — seeded from streaming, CWRU calibration pending. Envelope band/region reused verbatim from config/cwru.json (config/mfpt.json's _envelope_band_note) — zero new DSP tuning for this cross-rig claim.
> **Companion notes** (hand-authored, not regenerated by this runner): [`field_validation_part_c.md`](field_validation_part_c.md) — Live product-path (webapp) verification log — 5 real HTTP uploads drafted end-to-end through the agent path, hand-recorded after a live run.
## Part A — Rig files (known labels, frozen constants)
| File | Fault type | Load (lbs) | Expected | Detected | Confidence | Matched (Hz) | Computed (Hz) | Verdict |
|---|---|---|---|---|---|---|---|---|
| baseline_1 | normal | 270.0 | none | none | — | — | — | PASS |
| baseline_2 | normal | 270.0 | none | none | — | — | — | PASS |
| baseline_3 | normal | 270.0 | none | bearing_inner_race | high | 120.17 | 118.88 | FAIL (false positive) |
| inner_vload_1 | inner | 0.0 | bearing_inner_race | bearing_inner_race | high | 118.0 | 118.88 | PASS |
| inner_vload_2 | inner | 50.0 | bearing_inner_race | bearing_inner_race | high | 118.0 | 118.88 | PASS |
| inner_vload_3 | inner | 100.0 | bearing_inner_race | bearing_inner_race | high | 118.0 | 118.88 | PASS |
| inner_vload_4 | inner | 150.0 | bearing_inner_race | bearing_inner_race | high | 118.0 | 118.88 | PASS |
| inner_vload_5 | inner | 200.0 | bearing_inner_race | bearing_inner_race | high | 118.0 | 118.88 | PASS |
| inner_vload_6 | inner | 250.0 | bearing_inner_race | bearing_inner_race | high | 118.0 | 118.88 | PASS |
| inner_vload_7 | inner | 300.0 | bearing_inner_race | bearing_inner_race | high | 117.67 | 118.88 | PASS |
| outer_270_1 | outer | 270.0 | bearing_outer_race | bearing_outer_race | high | 80.5 | 81.12 | PASS |
| outer_270_2 | outer | 270.0 | bearing_outer_race | bearing_outer_race | high | 80.67 | 81.12 | PASS |
| outer_270_3 | outer | 270.0 | bearing_outer_race | bearing_outer_race | high | 80.67 | 81.12 | PASS |
| outer_vload_1 | outer | 25.0 | bearing_outer_race | bearing_outer_race | high | 80.0 | 81.12 | PASS |
| outer_vload_2 | outer | 50.0 | bearing_outer_race | bearing_outer_race | high | 80.33 | 81.12 | PASS |
| outer_vload_3 | outer | 100.0 | bearing_outer_race | bearing_outer_race | high | 80.33 | 81.12 | PASS |
| outer_vload_4 | outer | 150.0 | bearing_outer_race | bearing_outer_race | high | 80.33 | 81.12 | PASS |
| outer_vload_5 | outer | 200.0 | bearing_outer_race | bearing_outer_race | high | 80.33 | 81.12 | PASS |
| outer_vload_6 | outer | 250.0 | bearing_outer_race | bearing_outer_race | high | 80.67 | 81.12 | PASS |
| outer_vload_7 | outer | 300.0 | bearing_outer_race | bearing_outer_race | high | 80.67 | 81.12 | PASS |
### Part A pass bar
- Baseline clean (0 false bearing calls): NO (1 false positive(s)) — required: 100% (HARD LINE)
- IR/OR primary-diagnosis correct: 17/17 (100.0%) — required: ≥85%
- Confidence-vs-load health check: no variation in load or confidence across IR/OR rows — cannot compute a correlation
**Part A: MISS**
Per the frozen-constants invariant: a missed bar is reported here, not silently fixed by tuning config/thresholds.json or config/mfpt.json. Any proposed change requires a written diff + rationale for review.
## Part B — Real-world files (NOT officially labeled; scored per the two-outcome rule)
COMMITTED = a primary bearing-family finding matching one of the file's own embedded fault frequencies (or a harmonic/±1× shaft sideband), confidence ≥ medium. HONESTLY UNCERTAIN = low confidence (or no finding) with an explicit DIAGNOSIS-resolving follow-up recommended. Either is PASS. FAIL = a confident finding at an unrelated frequency, or a clean bill of health with no diagnosis-resolving follow-up. (The ISO-severity velocity follow-up now emitted on these acceleration-only files is a COVERAGE measurement -- it resolves the severity zone, not the fault -- and is deliberately excluded from this rule.)
### real_world_intermediate_speed_bearing — **FAIL** (confident diagnosis at frequencies unrelated to the embedded set)
- Shaft rate: 6.3289 Hz
- Embedded fault-frequency targets (Hz): ball=24.3, cage=2.76, outer=51.9, inner=67.85
- Detected: bearing_ball_spin (confidence: high)
- Recommended follow-up: Velocity measurement per ISO 20816 (10-1000 Hz broadband, mm/s RMS)
### real_world_oil_pump_bearing — **FAIL** (clean bill of health on a known-faulted machine, no diagnosis-resolving follow-up recommended)
- Shaft rate: 9.5155 Hz
- Embedded fault-frequency targets (Hz): ball=35.68, cage=7.9, outer=78.5, inner=114.19
- Detected: none (confidence: —)
- Recommended follow-up: Velocity measurement per ISO 20816 (10-1000 Hz broadband, mm/s RMS)
### real_world_planet_bearing — **FAIL** (clean bill of health on a known-faulted machine, no diagnosis-resolving follow-up recommended)
- Shaft rate: 0.6229 Hz
- Embedded fault-frequency targets (Hz): ball=2.95, cage=0.7, outer=6.38, inner=8.12
- Detected: none (confidence: —)
- Recommended follow-up: Velocity measurement per ISO 20816 (10-1000 Hz broadband, mm/s RMS)
**Part B: MISS (0/3 files PASS)**
## Overall: MISS
## Known Limitations
- **baseline_3 false positive (bearing_inner_race, high confidence)**: root-caused, not guessed. Its matched peak (120.17Hz, 1.09% from computed BPFI 118.88Hz) is the LOUDEST bin on its axis, but baseline_3's whole spectrum sits at a noise-floor scale -- max amplitude 0.0173, essentially the same order of magnitude as baseline_1/baseline_2's clean spectra (0.0151/0.0158) and ~8x SMALLER than a genuine fault peak on this same rig (outer_270_1's real BPFO peak: max amplitude 0.1395). _amplitude_is_marginal() (bearing_rca.py) only compares a peak to its OWN axis's ceiling, so the loudest bin of an otherwise-uniformly-quiet spectrum trivially clears that ratio and is never flagged marginal -- there is no absolute or cross-reading amplitude floor. PROPOSED (not applied): an absolute or fleet/baseline-referenced amplitude floor alongside the existing per-axis ratio check -- a bearing_rca.py change, flagged for review, not made here.
- **real_world_oil_pump_bearing / real_world_planet_bearing (FAIL, clean bill of health on known-faulted machines)**: root cause is the SAME mechanism already identified, attempted, and reverted during the Phase 3.5 CWRU OR014@6 investigation (see config/thresholds.json profiles.route.rca._rationale_spectrum_peak_selection) -- peaks_from_spectrum()'s top-3-peaks-per-axis-by-raw-amplitude selection (still at the uncalibrated defaults: cap=3, prominence=None) crowds out real bearing-fault-adjacent content when a louder, unrelated low-frequency peak dominates the spectrum. Pulled spectra confirm real candidate content exists but never enters the top-3: oil pump's top-3 peaks are at 0.05/7.15/140.45Hz, yet its true outer-race window (78.50Hz) has a local peak of 0.0017 and its inner-race window (114.19Hz) has a local peak of 0.0048 -- both above the spectrum's own mean amplitude (0.0009) but outranked by louder, unrelated content. Planet bearing's top-3 are 0.05/2.95/0.10Hz; its true BPFO/BPFI windows (6.38/8.12Hz) have local peaks of 0.0023/0.0010, again never in the top 3. This is real, independent cross-rig corroboration for the already-proposed-but-not-applied fix (a per-axis RELATIVE prominence, not a global absolute one) -- still not applied without a further review round, exactly per the existing rationale note. Compounding gap (partially addressed in Session A): recommend_measurements() (pdm_core/recommendations.py) still has no DIAGNOSIS-resolving rule keyed on 'gate passed, spectra present, zero bearing findings' -- the 6 diagnosis_rules (config/next_measurements.json) fire only on an EXISTING ambiguous/low-confidence finding, so a genuine clean-bill-of-health case still gets no fault-resolving follow-up even when weak-but-real candidate content, as pulled above, was sitting just outside the current peak-selection window. Session A DID add a severity-coverage follow-up (a velocity measurement) that now fires on these acceleration-only files, but it is a coverage measurement (resolves the ISO zone, not the fault) and is excluded from Part B scoring above. PROPOSED (not applied): a new diagnosis_rule + predicate for the clean-bill-with-marginal-candidate case -- flagged for review (Session B, detector scope).
- **real_world_intermediate_speed_bearing (FAIL, classified 'unrelated frequency' -- actually a near-miss)**: bearing_ball_spin WAS correctly found at high confidence, matched at 23.40Hz to the synthetic equivalent bearing's computed BSF (23.34Hz, 0.27% off) -- this peak is rank 3 on its axis, so peak SELECTION worked correctly here, unlike the two files above. The FAIL is an artifact of Part B's own scoring: the file's independently embedded ball order converts to 24.30Hz, a 3.72% delta from the detected 23.40Hz -- just outside route's 3% match tolerance. Root cause: equivalent_bearing_from_frequencies() derives geometry from the outer/inner orders ONLY (by design -- its own docstring states BSF/FTF are close approximations, not exact reproductions); for this file that approximation puts the synthetic BSF ~4% from the file's true ball frequency. The detector correctly matched its OWN computed (approximate) BSF -- this is a Part B eval-scoring precision limit meeting a documented approximation, not a pdm_core defect. PROPOSED (eval-scoring only, would not touch pdm_core or config): widen Part B's ball/cage-family match tolerance to account for the known geometry-approximation error -- not applied this run.
- **Bechhoefer companion findings (web search attempted, per Part B instructions)**: his 2016 paper "A Quick Introduction to Bearing Envelope Analysis" (hosted historically at mfpt.org, mirrored on Scribd) is the known source for MFPT's own commentary on these three real-world files, but the accessible copy truncates before the real-world case studies. Independent third-party dataset documentation (e.g. github.com/opprud/bearing_dataset's MFPT README) corroborates that these three files' fault locations are 'unknown' / not officially labeled by MFPT itself -- consistent with this phase's NO FABRICATED GROUND TRUTH rule. No specific per-file fault-location diagnosis from Bechhoefer's own analysis could be verified. TODO: revisit if a complete copy of the paper becomes accessible.
## Companion documents
Hand-authored sections live in their own files so a regeneration of this machine-written doc (`eval/runner.py`) never overwrites them:
- [`field_validation_part_c.md`](field_validation_part_c.md) — Live product-path (webapp) verification log — 5 real HTTP uploads drafted end-to-end through the agent path, hand-recorded after a live run.
← Back to the validation record