Executive Summary
A measurement without recognized meaning cannot support a compliance decision, and today no AI-derived substructure measurement carries recognized meaning in the federal track safety framework. The regulatory architecture measures track geometry, the surface expression of substructure condition, and the Federal Railroad Administration’s 2024 rulemaking on Track Geometry Measurement System inspections would codify automated inspection for geometry while listing ground penetrating radar and machine learning visual inspection among available technologies without assigning either a compliance role.
This paper examines the three problems that stand between the sensing capability established in Paper 1 and regulatory standing. The first is metric proliferation: the research community maintains multiple fouling quantification schemes, including the fouling index, percentage void contamination, the void contaminant index, and the relative ballast fouling ratio, and these schemes can classify the same ballast sample differently, which is disqualifying for a compliance metric. The second is validation design: FRA’s research discipline requires that new inspection methods demonstrate performance at least as good as the methods they supplement, yet substructure sensing measures conditions no human inspection method observes, so equivalence must be redefined against ground truth and outcomes rather than against a human baseline. The third is the threshold problem: a recognized metric implies a threshold, a threshold implies a remediation obligation, and GAO has documented that this chain operates as an adoption disincentive when newly visible marginal conditions create obligations that did not previously exist.
The geometry automation record, from the 2018 BNSF test program through the waiver, the advisory committee’s failure to reach consensus, litigation, and the 2024 rulemaking, supplies a tested institutional sequence and a set of cautionary lessons. The paper closes with a proposed five-step standing pathway for substructure metrics, running from internal maintenance use through industry recommended practice, structured test programs, conditional waivers, and rulemaking, with remediation criteria negotiated before fielding rather than after.
1. The Standing Problem
Paper 1 established what can be measured. This paper asks what those measurements are allowed to mean.
The distinction is institutional. A railroad can use a GPR-derived fouling estimate internally to schedule undercutting today; nothing prevents it. What the estimate cannot do is substitute for, supplement, or satisfy any requirement of the federal track safety standards, because those standards, at 49 CFR Part 213, define track condition in terms of geometry parameters, visual inspection findings, and rail integrity, and no substructure condition index appears among them. FRA’s 2024 rulemaking would extend the framework by requiring qualifying Track Geometry Measurement System (TGMS) inspections at specified frequencies; in that document the agency catalogs ground penetrating radar, machine learning based visual inspection of track components, vertical track deflection systems, and lidar scanning as technologies now used to measure track health, while assigning compliance standing only to geometry measurement.
The consequence, identified as Finding 2 of the foundational paper, is that substructure AI generates value only through internal maintenance planning, which weakens the investment case GAO identified as decisive: railroads will not adopt a technology without confidence in a positive return. Standing is therefore an economic question as much as a legal one, and the route to standing runs through two technical prerequisites, a defensible metric and a defensible validation, before it reaches the institutional sequence examined in Section 4.
2. The Metric Landscape
2.1 Sample-based fouling quantification
The foundational quantity is the fouling index (FI) of Selig and Waters, defined from sieve analysis as the sum of the percentage of material passing the 4.75 mm sieve and the percentage passing the 0.075 mm sieve, with classification bands running from clean (FI below 1) through moderately clean (1 to below 10), moderately fouled (10 to below 20), fouled (20 to below 40), and highly fouled (40 and above). The index has competitors. Percentage void contamination (PVC) measures the fraction of ballast void space occupied by fouling material; the void contaminant index (VCI) refines the void-based approach; and the relative ballast fouling ratio was proposed specifically because the earlier parameters have limitations, including insensitivity to the density and character of the fouling material.
Two properties of this landscape matter for standardization. First, the schemes disagree: field studies have documented samples classified as moderately clean under the fouling index and moderately fouled under percentage-of-fouling measures, with the study authors concluding that a comprehensive treatment of fouling scales and classification is needed. A metric intended to carry compliance weight cannot classify the same physical condition into different categories depending on the formula chosen. Second, the schemes are gradation-based and therefore blind to what the gradation is made of; fouling material character matters to performance, a limitation the relative ratio approach was designed to address.
2.2 The performance connection
A compliance metric must correlate with outcomes that safety regulation cares about. The engineering literature supplies that connection for fouling: settlement behavior after tamping varies systematically with fouling category and moisture, with fouled, wet ballast settling more and producing greater geometry roughness under traffic, based on test facility and laboratory results spanning the classification bands. Reviews of fouled ballast definitions also draw a distinction the standards discussion needs: most existing limits worldwide are maintenance limits set by individual railroads for their own planning purposes rather than safety limits, and the two serve different functions and carry different evidentiary burdens.
2.3 Sensor-derived indices
Paper 1 documented the measurement side of this landscape: GPR-derived condition scores integrating dielectric permittivity, thickness scatter, boundary signal strength, and frequency spectrum area, verified against field fouling data; machine learning prediction of the fouling index directly from radar signal parameters; and deep learning image assessment of ballast sections against sieve-derived ground truth. The structural point for standardization is that every sensor-derived index is calibrated against a sample-based metric, so the ambiguity documented in Section 2.1 propagates: a GPR index trained against the fouling index inherits the fouling index’s disagreements with the void-based schemes. Metric selection therefore precedes, and constrains, sensor validation.
3. The Validation Problem
3.1 The equivalence discipline and its limits
FRA’s research program maintains a standing discipline for inspection innovation: methods for checking that new inspection systems are at least as good as those they supplement or replace. For geometry automation the discipline is directly applicable, because automated and visual inspection observe overlapping defect sets, and the comparison can be run head to head; the BNSF test program did exactly that, and FRA found that for every geometry defect identified by visual inspection, the automated system identified over 200.
Substructure sensing breaks the comparison’s premise. Visual inspection observes surface symptoms of substructure distress; it does not measure fouling percentage, layer thickness, trapped moisture, or track modulus at all. A GPR fouling estimate has no human-inspection baseline to be at least as good as. Demanding head-to-head equivalence would therefore either be trivially satisfied (any measurement exceeds no measurement) or incoherently framed. Validation must be redesigned around what the measurement claims.
3.2 Three validation constructs
The evidence assembled across Papers 1 and 2 supports a three-construct validation architecture, ordered by directness.
Ground-truth-referenced validation compares the sensor-derived index against the physical measurement it estimates: GPR fouling indices against sieve analysis of excavated samples, image-based condition scores against laboratory gradation, deflection-derived subgrade condition against geotechnical investigation. This construct is already exercised in the literature, including field verification of GPR condition scoring and sieve-referenced image assessment. Its cost is the cost of ground truth, established in Paper 1 as the root of the data scarcity problem.
Performance-referenced validation tests whether the index predicts the engineering behavior it is supposed to predict: whether locations scored as highly fouled exhibit the elevated settlement and geometry roughness the degradation literature associates with that condition. This construct connects the metric to consequences without requiring excavation at every validation point.
Outcome-referenced validation operates at program scale: whether inspection and maintenance programs informed by the index produce measurably better track condition trends than programs without it. The geometry record again supplies the template; the industry’s waiver filings documented defect-ratio improvement over the life of the test programs on covered territory. Outcome validation is the slowest construct and the one regulators ultimately credit.
A standing pathway should require all three in sequence, and Paper 5’s assurance pillar will incorporate this architecture, adding in-service performance monitoring against the drift hazards Paper 4 catalogs.
4. The Geometry Precedent
The automated geometry inspection record is the only completed instance of an inspection technology moving from research capability toward regulatory standing in the modern track safety framework, and every stage of it is documented in the public record.
Test program. In 2018, FRA approved a BNSF test program specifically designed to evaluate the effectiveness of automated track inspection technologies, granting a temporary, limited suspension of the visual inspection frequency requirements of 49 CFR 213.233(c) as necessary to conduct it. The design point matters: FRA explained in related litigation that continuing the full manual inspection schedule would have prevented the program from determining whether a specific combination of visual and automated inspections produces the greatest results for safety and operations. A test program that changes nothing measures nothing.
Waiver. In 2021, BNSF concluded the test program and FRA approved a waiver allowing continued use of the test methodologies on designated track with additional safety metrics in place. The waiver record also fixed remediation practice: identified defects were verified centrally, with a 24-hour window for lower-severity classifications and immediate slow orders for higher-severity ones.
Advisory process. FRA tasked a Railroad Safety Advisory Committee working group with the combined visual and automated inspection question. Every test program had concluded by November 2022; in October 2023 the working group determined it would not be able to reach consensus, while agreeing that automated inspection technology benefits track safety, and the task was closed in March 2024 without a recommendation. [CONTESTED TERRAIN] The underlying dispute, whether automated inspection should permit reduced visual inspection frequencies or should only ever supplement them, remained unresolved through litigation between BNSF and FRA over waiver expansion and through subsequent waiver proceedings, with the maintenance-of-way labor organization opposing frequency reductions and industry petitioners seeking them. This series carries that dispute as contested terrain; nothing in this paper resolves it.
Rulemaking. FRA then proceeded on its own record, proposing in October 2024 to require qualifying TGMS inspections at specified frequencies, basing the proposal in part on its research, ATIP operational experience, the test program results, and the BNSF waiver. Notably, the proposal codified automated inspection requirements while leaving the visual inspection regulations untouched, and its proposed one-hour remediation timeframe departed from the industry practice range documented in the waiver record.
Lessons for substructure metrics. Four transfer directly. The sequence works: test program, waiver, advisory engagement, rulemaking is a viable institutional ladder even when the advisory step fails to converge. The test program must be permitted to vary practice, or it cannot generate evidence. The remediation protocol is negotiated at the waiver stage, and its terms (verification workflow, response windows, severity tiers) become the de facto template the eventual rule reacts to. And the contested labor and oversight question does not resolve itself along the way; a substructure program that waits for consensus on the automation dispute will wait indefinitely, while one that positions substructure sensing as additive capability, measuring what visual inspection cannot see at all, engages the dispute on materially different terms than geometry automation did, a positioning advantage the foundational paper identified.
5. Metric, Threshold, Obligation
The chain that gives a metric force is also the chain that deters its adoption. GAO documented the mechanism: stakeholders reported a disincentive under current regulations to use new track inspection technologies, because such technologies identify defects perceived as too insignificant to pose a safety risk, yet once identified, those defects require remedial action. For substructure sensing the mechanism operates prospectively: the moment a fouling threshold acquires regulatory meaning, every mile of newly measured track becomes potentially obligation-bearing, and the technology that reveals the condition becomes the technology that creates the liability.
Three design principles follow from the record. First, the maintenance-limit and safety-limit distinction documented in the fouled ballast literature should be preserved in any standard: maintenance thresholds trigger planning, safety thresholds trigger operational restriction, and collapsing them converts every planning signal into an operating restriction, which is precisely the disincentive GAO described. Second, remediation criteria for AI-detected substructure conditions should be negotiated at the test program and waiver stages, before fielding, following the geometry precedent in which verification workflow and response windows were fixed in the waiver terms. Third, threshold values should be set from the performance-referenced evidence base (settlement and geometry consequence data by fouling category and moisture state), not from measurement convenience, so that the obligation attaches to demonstrated consequence rather than to detectability.
6. A Standing Pathway for Substructure Metrics
Assembling Sections 2 through 5, the paper proposes a five-step pathway.
Step 1: Metric consolidation. Industry and research stakeholders converge on a single fouling quantification scheme and classification banding for standards use, resolving the documented inconsistencies among the fouling index, void-based, and ratio-based schemes, and specifying moisture state alongside gradation given its documented effect on performance. Recommended practice bodies are the natural venue; a recommended practice is standing short of regulation, and it gives sensor developers a fixed calibration target.
Step 2: Ground-truth-referenced validation. Sensor-derived indices are validated against the consolidated metric across the range of ballast materials, climates, and traffic profiles in service, with the validation data contributed to the shared corpus architecture Paper 1 outlined, so validation and training data accumulate together.
Step 3: Performance-referenced validation. Index thresholds are tied to demonstrated settlement and geometry consequences, establishing the evidentiary basis for the maintenance-limit tier and, where justified, a safety-limit tier.
Step 4: Structured test programs. Railroads operate substructure-informed maintenance programs under FRA-approved test programs that permit practice to vary, generating the outcome-referenced evidence, with remediation criteria for detected conditions defined in the program terms.
Step 5: Conditional waivers and rulemaking. Successful test programs convert to conditional waivers, and the accumulated record supports rulemaking that assigns substructure indices a defined role, whether as a qualifying input to risk-based inspection frequencies or as a recognized condition measure in maintenance standards. The geometry record demonstrates the full ladder is climbable within roughly a six-year span from test program approval to proposed rule, and also demonstrates that the advisory step may fail to converge without halting the sequence.
7. Findings
Finding 2.1: Metric proliferation is a standardization blocker in its own right. Multiple fouling quantification schemes coexist, they can classify identical samples differently, and every sensor-derived index inherits its calibration metric’s ambiguities; consolidation on a single scheme with moisture specification is the necessary first step, prior to any sensor question.
Finding 2.2: The equivalence discipline requires redesign for measurements without a human baseline. Head-to-head comparison against visual inspection, the construct that carried geometry automation, is unavailable for substructure sensing; a three-construct architecture of ground-truth-referenced, performance-referenced, and outcome-referenced validation is the defensible substitute.
Finding 2.3: The geometry record establishes the institutional ladder and its costs. Test program, waiver, advisory engagement, and rulemaking form a demonstrated sequence; the sequence survives advisory non-consensus; and the remediation protocol negotiated at the waiver stage becomes the template the rule reacts to.
Finding 2.4: The threshold-obligation chain is the central adoption risk and must be engineered deliberately. The GAO-documented disincentive operates with amplified force for a technology that newly reveals subsurface conditions; preserving the maintenance-limit and safety-limit distinction and fixing remediation criteria before fielding are the available controls.
Finding 2.5: Substructure sensing’s additive character changes its regulatory politics. [CONTESTED TERRAIN adjacent] Because it measures what visual inspection cannot observe, substructure sensing need not be framed as a substitute for inspector work, and the standing pathway can proceed without first resolving the automation and visual inspection dispute, though that dispute will shape any step that touches inspection frequencies.
8. Limitations
The regulatory events of late 2025, including reported judicial direction to FRA on waiver expansion and subsequent Railroad Safety Board action on an industry-wide waiver, are documented here from trade press accounts quoting the underlying letters and dockets; the primary documents should be pulled from the dockets and verified before this paper is used in any formal proceeding. Threshold values are deliberately not proposed in this paper; setting them requires the performance-referenced evidence assembly described in Step 3, which exceeds this installment’s scope. International standards regimes, including European fouling and maintenance limits, were noted in the definitional literature but not comparatively analyzed; that comparison is assigned to the post-series research agenda.
9. Threads Passed Forward
The economic logic implicit in Section 5, that standing converts substructure measurement from an internal planning tool into a compliance-relevant asset and thereby changes its return profile, is the opening premise of Paper 3, which constructs the value case in full. The threshold-obligation chain and the possibility of metric gaming or threshold-adjacent behavior are carried to Paper 4 as institutional hazards, alongside the validation-decay risk that arises when a validated model meets drifting field conditions. The five-step pathway of Section 6 becomes the regulatory engagement pillar of Paper 5’s framework, where it is integrated with the assurance, data governance, teaming, and economic pillars and assigned phase gates.
Bibliography
Tier 1: Government and Oversight Sources
Congressional Research Service. Freight Rail Safety Issues in the 119th Congress. CRS Report R47911. Washington, DC. https://www.congress.gov/crs-product/R47911.
Federal Railroad Administration. FRA Research, Development and Technology Strategic Plan, 2020-2024. Washington, DC: U.S. Department of Transportation, 2020. https://railroads.dot.gov/sites/fra.dot.gov/files/2020-07/Strategic%20Plan%202020-2024-A.pdf.
Federal Railroad Administration. “Petition for Waiver of Compliance.” 83 Fed. Reg. 55450 (November 5, 2018). Docket No. FRA-2018-0091.
Federal Railroad Administration. “Track Geometry Measurement System (TGMS) Inspections.” Notice of Proposed Rulemaking. Federal Register, October 24, 2024. Docket No. FRA-2024-0032. https://www.federalregister.gov/documents/2024/10/24/2024-24153/track-geometry-measurement-system-tgms-inspections.
Federal Railroad Administration. BNSF Waiver Letter and associated decision documents. Docket No. FRA-2020-0064, Regulations.gov.
U.S. Government Accountability Office. Rail Safety: Federal Railroad Administration Should Report on Risks to the Successful Implementation of Mandated Safety Technology. GAO-11-133. Washington, DC, 2010. https://www.gao.gov/products/gao-11-133.
Docket Filings and Proceedings
Association of American Railroads and American Short Line and Regional Railroad Association. Comments on Track Geometry Measurement System (TGMS) Inspections NPRM. Docket No. FRA-2024-0032, January 2025.
Tier 2 and Tier 4: National Academies, FRA-Sponsored, and Peer-Reviewed Sources
Chrismer, Steven, and James Hyslip. “Principles of Degraded Ballast and Their Track Safety Implications.” Technical paper drawing on Transportation Technology Center test results.
“Development of Condition Assessment Index of Ballast Track Using Ground-Penetrating Radar (GPR).” Applied Sciences (2021). https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8539047/.
Fouled Ballast Definitions and Parameters. FRA-sponsored technical report, 2017.
Indraratna, Buddhima, et al. “A New Parameter for Classification and Evaluation of Railway Ballast Fouling.” Canadian Geotechnical Journal 48 (2011): 322-326.
Luo, Jiayi, et al. “Toward Automated Field Ballast Condition Evaluation: Development of a Ballast Scanning Vehicle.” Transportation Research Record (2024). https://doi.org/10.1177/03611981231178302.
Selig, Ernest T., and John M. Waters. Track Geotechnology and Substructure Management. London: Thomas Telford, 1994.
“Study of Ballast Fouling in Railway Track Formations.” Peer-reviewed field study of fouling classification consistency (2012).
Tier 5: Trade Press (attributed context only)
Railway Age and Trains reporting on FRA waiver proceedings, litigation, and Railroad Safety Board actions, 2022-2025, cited solely for the timeline of proceedings pending primary-document verification.
A PDF of this paper is coming
The full text is on this page and nothing is held back. Leave your address and we will send the PDF as soon as it is rendered.