Experimental Typst web edition · Veto chapter pilot

5 Shielding and Veto System

Introduction

The current BabyIAXO working baseline assumes an on-surface installation at DESY, where passive shielding must suppress environmental photons without the benefit of underground overburden. The same lead that provides this attenuation does not efficiently moderate fast neutrons and can convert them into secondary neutrons, photons, and charged fragments. The active shield must therefore perform two functions: reject through-going cosmic muons and tag the distributed prompt and post-trigger activity associated with neutron-induced showers.

This chapter presents the physical argument and experimental evidence that led to a three-stage scintillator–cadmium veto surrounding the lead shield. The narrative is organized around the design decisions: why passive shielding is insufficient, how the active stages sample the neutron-induced shower, why three layers were selected, how the waveform information is reduced to veto observables, and what the commissioned IAXO-D0 prototype demonstrates in data. Detailed nuclear-data benchmarks, campaign metadata, full diagnostic scans, and the technical construction of the late-window control analysis are retained in the Appendix.

The division of scope with Section 6 is explicit. This chapter addresses veto physics, conditional rejection, calibration-like control acceptance, and prototype validation, including the published IAXO-D0 data cut flow. Section 6 provides the source normalization, confidence intervals on source-normalized simulation levels, and prompt/delayed accounting used in the conservative partial source-response model. The common REST-for-Physics/restG4 implementation is described in the software [8386].

5.1 Design boundary and source inputs

5.1.1 Detector configurations and requirements

Three related detector configurations appear in this chapter and should not be conflated. The optimization simulations use the 59-panel veto geometry developed for the detector line. The experimental validation uses the commissioned 57-panel IAXO-D0 surface prototype. A detector-and-site prediction for BabyIAXO at DESY can be made only after the corresponding contracts are frozen; it is not inferred by rescaling the Zaragoza prototype result.

ConfigurationRole in this chapterQuantitative use
59-panel simulated vetoThree scintillator–cadmium stages around the lead shieldGeometry and layer optimization; conditional waveform-level rejection
57-panel IAXO-D0 prototypeCommissioned surface setup in ZaragozaPublished 52.1-day data cut flow and waveform-level validation
BabyIAXO projection at DESYPlanned detector and site configuration; not yet frozenRequires site-specific source transfer, final channel map, and a separately validated response contract
Table 5.1: Detector scopes used in the shielding and veto discussion. Numerical results are not transferred between configurations without stating the corresponding geometry, source, and response definition.

The design must preserve X-ray-like Micromegas events in the keV region while rejecting surface backgrounds. It must remain compatible with the X-ray beam path, services, mechanical clearances, available scintillator modules, and the synchronized AGET-based readout. The beam pipe prevents perfectly hermetic passive shielding, whereas the neutron-capture timescale requires a readout window long enough to retain post-trigger activity without accepting an excessive number of random coincidences. These constraints make geometry, timing, and analysis inseparable parts of the veto design.

5.1.2 Surface source inputs

The selected design studies use source-specific inputs rather than one generic “cosmic” sample. Production muons use the correlated sea-level parameterization of Guan et al. implemented by the CosmicMuons generator [104]. The neutron layer scans use the outdoor HENSA spectrum in the 1MeV10GeV interval [105, 106]. Their nominal input histogram is proportional to sin𝜃cos2𝜃, but the inspected projected-source reader applies an additional 1/cos𝜃 factor, producing a sampled density proportional to sin𝜃cos𝜃. The input histogram and sampled angular law must therefore be distinguished when interpreting the layer and inclination scans. The campaign normalization and angular-measure audit are documented in Section 6; the geometry comparisons here retain their common historical source setting. Historical passive-shield scans and cross-checks use CRY, while EXPACS provides an independent spectral comparison [88, 107, 108].

The HENSA outdoor field includes atmospheric and environment-modified neutrons measured at the surface. It is therefore described below as the HENSA outdoor neutron field, not as a pure cosmic component. It must not be added to the HENSA-minus-CRY environmental residual as if the two were independent source terms. The generator implementation is documented in Section 4.6.5; the source normalization and non-additivity convention are defined in Section 6.5.1.4.

5.2 Limits of passive shielding

5.2.1 Material roles

Lead remains essential because it strongly attenuates environmental photons. Its role for fast neutrons is different. In an elastic collision with a nucleus of mass number 𝐴, the maximum transferable fraction of neutron kinetic energy is 4𝐴/(1+𝐴)2, which is unity for hydrogen but only about 0.019 for lead. Consequently, lead is a poor moderator, while inelastic and neutron-emission channels produce a secondary shower once the incident energy reaches the MeV range.

Hydrogen-rich plastic scintillator complements the lead by slowing secondary neutrons and converting recoil-proton and charged-particle energy into light. Cadmium is placed only after this moderation step: its large thermal-neutron capture probability converts the slowed component into a gamma cascade near an active panel. The nuclear de-excitation is prompt relative to the capture, but moderation can place the capture and its scintillator response well after the Micromegas trigger. The signal medium itself remains sensitive to compact secondary photons and electrons as well as to nuclear recoils, so the veto must tag the surrounding history rather than identify the primary particle from the Micromegas deposit alone.

The transport-model checks that support this mechanism are summarized in Section 5.3.3 and documented in Appendix Section A.2.

5.2.2 Lead thickness and penetrations

A parameterized lead-shield scan was performed with idealized 4𝜋 coverage and with the X-ray beam-pipe opening. The historical CRY samples are used here only to compare geometries under a common source definition; they do not define the final absolute background level. The full per-particle results are retained in Appendix Section A.11.

Figure 5.1: Lead-thickness scan for the two components that determine the passive-shielding decision. Photon-induced activity is strongly attenuated, although leakage through the beam-pipe opening limits the ideal suppression. Neutron-induced activity remains weakly dependent and non-monotonic because fast-neutron interactions in lead generate secondary particles. The ordinate is the historical scan response and is used comparatively, not as the current background-model normalization.
Figure 5.1: Lead-thickness scan for the two components that determine the passive-shielding decision. Photon-induced activity is strongly attenuated, although leakage through the beam-pipe opening limits the ideal suppression. Neutron-induced activity remains weakly dependent and non-monotonic because fast-neutron interactions in lead generate secondary particles. The ordinate is the historical scan response and is used comparatively, not as the current background-model normalization.

The photon response improves by more than four orders of magnitude over the idealized scan and by roughly a factor 3×102 when the pipe opening is included. The neutron response remains within the same order of magnitude and exhibits a broad intermediate-thickness enhancement. The design consequence is robust: thick lead is retained for photon suppression, but increasing its thickness does not replace an active neutron-sensitive system.

5.2.3 Passive neutron moderation

Borated high-density polyethylene (HDPE) was tested as a moderator and absorber while retaining the 20cm total lead thickness required for photon attenuation [105]. The scan varied both the HDPE thickness and its position within a Pb/HDPE/Pb stack. Hydrogen slows the neutron population, and the dominant excited-state branch of 10B(n,𝛼)7Li releases a 477.6keV gamma after capture. The usefulness of this mechanism depends on where moderation and capture occur relative to the lead and the detector.

(a) Scanned layer ordering.
(a) Scanned layer ordering.
(b) Residual neutron response.
(b) Residual neutron response.
Figure 5.2: Passive Pb/borated-HDPE/Pb study with fixed total lead thickness. The best sampled configuration reduces the neutron-induced response by about 40%, but the gain depends strongly on material ordering and is insufficient as a stand-alone surface-neutron mitigation.

The most favorable configurations reduce the neutron response by only 3040%, corresponding to a suppression factor of order 1.6. Practical moderator thicknesses do not thermalize the full MeV–GeV component before secondary production has occurred in the lead and surrounding structures. This result defines the active-veto boundary condition: the lead remains in place, and the active panels are arranged to sample the shower that it produces.

5.3 Neutron-sensitive active veto concept

5.3.1 From surface primaries to taggable secondaries

The integral surface-neutron field is not, by itself, the population that controls the detector response. Low-energy neutrons are abundant in the outdoor spectrum, but the primaries that produce Micromegas-sensitive activity after traversing the shielding are much harder. In the HENSA three-layer diagnostic, 59.68% of the detector-reaching population lies between 100MeV and 1GeV, and a further 31.78% lies between 1 and 10GeV. The historical CRY cross-check exhibits the same qualitative hardening. The comparison in Figure 5.3 is normalized within each source and detector-reaching sample; it establishes the relevant energy range but is not an absolute source comparison.

Figure 5.3: Primary-neutron energy distributions before and after requiring Micromegas-sensitive activity. The upper panel compares the normalized CRY and HENSA outdoor source shapes; the lower panel shows the corresponding detector-reaching populations. The HENSA production is restricted to 1 MeV – 10 GeV . The hardening demonstrates that the detector-relevant population is selected from the high-energy tail rather than from the dominant low-energy source fluence.
Figure 5.3: Primary-neutron energy distributions before and after requiring Micromegas-sensitive activity. The upper panel compares the normalized CRY and HENSA outdoor source shapes; the lower panel shows the corresponding detector-reaching populations. The HENSA production is restricted to 1MeV10GeV. The hardening demonstrates that the detector-relevant population is selected from the high-energy tail rather than from the dominant low-energy source fluence.

Once these fast primaries interact, the lead converts them into a much softer and more numerous secondary population. In the studied 1745 detector-reaching HENSA events, approximately 9.19×104 neutrons are produced in lead, 6.58×104 leave the lead volume, and 1.83×104 leave with momentum directed toward the TPC. Their energy distribution peaks around the MeV scale, where moderation in hydrogen-rich scintillator and subsequent capture become effective. The directionality also shows that the shower is distributed throughout the shield rather than confined to the detector-facing surface.

Figure 5.4: Secondary neutrons produced in the lead shielding in the HENSA outdoor three-layer simulation. The upper panel shows the production and escape spectra as 𝐸 𝑑 𝑁 / 𝑑 𝐸 per detector-reaching event. The lower panel gives the fraction of exiting neutrons directed toward the TPC. Lead shifts the relevant population toward the MeV scale and produces a distributed shower that can be sampled by the surrounding active stages.
Figure 5.4: Secondary neutrons produced in the lead shielding in the HENSA outdoor three-layer simulation. The upper panel shows the production and escape spectra as 𝐸𝑑𝑁/𝑑𝐸 per detector-reaching event. The lower panel gives the fraction of exiting neutrons directed toward the TPC. Lead shifts the relevant population toward the MeV scale and produces a distributed shower that can be sampled by the surrounding active stages.

5.3.2 Tagging the secondary shower

The central design principle is to tag the neutron-induced cascade rather than the incoming neutron. Fast primaries interact in the lead and copper, producing softer neutrons, photons, recoil nuclei, and charged fragments. Plastic scintillator samples the prompt charged component and moderates part of the neutron population. Cadmium sheets between active stages then capture a fraction of the thermalized component and emit gamma cascades close to a neighboring scintillator.

Figure 5.5: Physical origin and timing of the two neutron-sensitive veto signatures. A fast surface neutron initiates a secondary cascade in the lead shielding. Charged secondaries and proton recoils produce near-trigger scintillation, whereas moderated neutrons may be captured in cadmium; the resulting gamma cascade deposits energy in neighboring plastic. The capture gamma emission is prompt relative to the nuclear capture, but the preceding moderation places many such responses after the Micromegas trigger. The lower panel uses this trigger-referenced convention; the dashed curve denotes qualitative capture-time density, not pulse amplitude. Geometry is schematic.
Figure 5.5: Physical origin and timing of the two neutron-sensitive veto signatures. A fast surface neutron initiates a secondary cascade in the lead shielding. Charged secondaries and proton recoils produce near-trigger scintillation, whereas moderated neutrons may be captured in cadmium; the resulting gamma cascade deposits energy in neighboring plastic. The capture gamma emission is prompt relative to the nuclear capture, but the preceding moderation places many such responses after the Micromegas trigger. The lower panel uses this trigger-referenced convention; the dashed curve denotes qualitative capture-time density, not pulse amplitude. Geometry is schematic.

The need to tag this surrounding history is visible directly in the saved-event truth record. In a sample of 879 neutron-induced events with a TPC energy deposit, electromagnetic descendants account for 552 events (62.8%), gas recoils or nuclear fragments for 242 (27.5%), charged-meson descendants for 41 (4.7%), and direct neutron interactions in the gas for only 44 (5.0%). Thus, in almost two thirds of this sample the low-energy TPC activity is produced through the electromagnetic branch of a wider neutron-induced cascade. The classification is used only to establish the mechanism; it is neither an experimental particle label nor a source-normalization factor.

Figure 5.6: Dominant route to the TPC energy deposition in the 879-event neutron-history sample, which contains a total of 43.2 MeV deposited gas energy. Solid bars give event fractions with central 68.27% Wilson intervals, while hatched bars give the corresponding descriptive fractions of total gas energy; no uncertainty is assigned to the latter because only aggregate energy summaries are available. Electromagnetic descendants dominate the event count, while recoils and fragments dominate the deposited energy; direct neutron interactions account for only 5.0% of events and 0.3% of the gas energy.
Figure 5.6: Dominant route to the TPC energy deposition in the 879-event neutron-history sample, which contains a total of 43.2MeV deposited gas energy. Solid bars give event fractions with central 68.27% Wilson intervals, while hatched bars give the corresponding descriptive fractions of total gas energy; no uncertainty is assigned to the latter because only aggregate energy summaries are available. Electromagnetic descendants dominate the event count, while recoils and fragments dominate the deposited energy; direct neutron interactions account for only 5.0% of events and 0.3% of the gas energy.

Simplified sandwich scans were used to establish where the active material should sit relative to the lead. The response was evaluated after Birks quenching in the scintillator, so the threshold represents an electron-equivalent visible-energy proxy rather than unquenched deposited energy [109, 110]. The full material-ordering and quenching diagnostics are preserved in Appendix Section A.7.

Figure 5.7: Representative simplified sandwich comparison for high-energy neutrons incident on a lead-coupled scintillator slab. The response is shown after Birks quenching for a PVT-like baseline, cadmium-assisted PVT, and boron-loaded scintillator. The scan isolates material ordering before light attenuation, waveform formation, and peak reconstruction are introduced.
Figure 5.7: Representative simplified sandwich comparison for high-energy neutrons incident on a lead-coupled scintillator slab. The response is shown after Birks quenching for a PVT-like baseline, cadmium-assisted PVT, and boron-loaded scintillator. The scan isolates material ordering before light attenuation, waveform formation, and peak reconstruction are introduced.

The bare scintillator responds most readily to neutrons that can interact directly in the active material. For the harder primaries that dominate Micromegas-triggering events, lead-coupled arrangements retain the useful response because they sample the secondary cascade. This is why the veto surrounds the passive shield instead of replacing it.

5.3.3 Transport-validation envelope

The two transport steps most directly connected to the design conclusion were checked independently. First, a thin-target benchmark of outgoing neutrons from natural lead at 14.1MeV was compared with EXFOR data in the measured energy and angular acceptance. The four tested high-precision Geant4 reference lists lie 89% below the experimental integral and differ from one another by less than 1% in this low-energy benchmark. This agreement supports the secondary-shower interpretation, although it does not by itself validate the full GeV-scale HENSA response.

Second, alternative cadmium de-excitation treatments were replayed through the same detector-response and peak-finding chain. The probability of reconstructing any veto peak remains between 0.719 and 0.742, whereas the conditional fraction above the 10MeV-equivalent operating point changes from 0.314 to 0.383. The existence of a neutron-sensitive signature is therefore more stable than its efficiency at a fixed reconstructed-energy threshold. The full cross-section comparisons, outgoing-neutron spectra, cadmium line yields, and response tables are retained in Appendix Section A.2; the threshold dependence enters the systematic envelope rather than being tuned to data.

5.3.4 Prompt and post-trigger signatures

Here, prompt and post-trigger refer to time relative to the Micromegas trigger, not to the duration of the nuclear de-excitation. Figure 5.8 connects the simulated capture history to reconstructed veto activity.

Figure 5.8: Capture truth and reconstructed response in 2,706 processed events from the HENSA outdoor three-layer sample; no conservative-reference Micromegas selection is applied. The upper panels identify the first-capture material per event and the material of all 73,295 explicit capture records; the latter is a descriptive record fraction because an event can contribute several correlated captures. The lower-left panel shows the cadmium-capture time relative to the first simulated gas deposit for the 2,453 events with prompt TPC activity; its band is a central 68.27% bootstrap interval obtained by resampling complete events. The event-fraction error bars in the upper-left and lower-right panels are central 68.27% Wilson intervals. The lower-right panel gives the event-level correspondence between cadmium capture and reconstructed veto peaks.
Figure 5.8: Capture truth and reconstructed response in 2,706 processed events from the HENSA outdoor three-layer sample; no conservative-reference Micromegas selection is applied. The upper panels identify the first-capture material per event and the material of all 73,295 explicit capture records; the latter is a descriptive record fraction because an event can contribute several correlated captures. The lower-left panel shows the cadmium-capture time relative to the first simulated gas deposit for the 2,453 events with prompt TPC activity; its band is a central 68.27% bootstrap interval obtained by resampling complete events. The event-fraction error bars in the upper-left and lower-right panels are central 68.27% Wilson intervals. The lower-right panel gives the event-level correspondence between cadmium capture and reconstructed veto peaks.

In the processed three-layer HENSA sample, 98.3% of analysis entries contain a cadmium-sheet capture and 78.9% of all explicit captures occur in cadmium. At event level, however, 23.8% contain a cadmium capture but no reconstructed veto peak. Approximately one quarter contain only prompt or near-trigger activity, while 49.9% contain a late reconstructed peak. Finite light collection, thresholds, timing, and reconstruction therefore prevent a one-to-one mapping between capture and observed veto signature.

The cadmium capture history in the HENSA three-layer simulation provides the characteristic post-trigger timescale relative to the first simulated gas deposit, which is used as a trigger proxy and is distinct from both the physical capture time and the reconstructed veto-peak time. A truncated-exponential fit to the selected post-trigger physical-capture tail gives 𝜏=46.9𝜇s. This is a simulation-derived capture-retention study, not a complete acquisition-window optimization: reconstructed thresholds, accidental activity, and dead time are not included in the fitted curve.

Figure 5.9: Simulation-derived cadmium-capture timing and conditional retention for the HENSA three-layer sample, measured from the first simulated gas deposit used as a trigger proxy. Of the selected positive-delay captures, 79.7 % occur by 70 mus and 88.8 % by 100 mus . The published prototype record spans approximately [ − 30 , + 70 ] 𝜇 s relative to its hardware trigger; mapping that trigger to the gas-deposit proxy is a separate response requirement. These truth-history fractions are not reconstructed veto efficiencies. The denominator and fit interval are documented in Appendix Section A.4 .
Figure 5.9: Simulation-derived cadmium-capture timing and conditional retention for the HENSA three-layer sample, measured from the first simulated gas deposit used as a trigger proxy. Of the selected positive-delay captures, 79.7% occur by 70mus and 88.8% by 100mus. The published prototype record spans approximately [30,+70]𝜇s relative to its hardware trigger; mapping that trigger to the gas-deposit proxy is a separate response requirement. These truth-history fractions are not reconstructed veto efficiencies. The denominator and fit interval are documented in Appendix Section A.4.

Late-window activity alone is not neutron-specific because calibration data contain electronic noise and unrelated environmental coincidences throughout the waveform. The useful observable combines prompt or near-trigger activity with post-trigger energy, peak multiplicity, signal quality, and segmentation. An event-mixing study quantifies the timing information: a late peak alone accepts approximately 52% of the neutron sample and 56% of a neutron-derived time-scrambled control, whereas requiring both prompt and late activity accepts approximately 46% and 14%, respectively. The scrambled control preserves neutron-like peak multiplicity and amplitude while removing alignment to the trigger; it does not measure the apparatus’s accidental-coincidence rate. The improvement comes from the timing context rather than from late activity by itself; the capture-denominator bookkeeping, full trade-off plot, and selection definitions are documented in Appendix Section A.4. In this chapter, post-trigger or late-window activity always refers to peaks inside the recorded waveform. The term delayed activation is reserved for radioactive-decay descendants occurring beyond the coincidence window, as defined in Section 6.7. No standalone 𝑚3 peak-multiplicity requirement is adopted as an operating point in this thesis. Multiplicity enters the structured and multivariate selectors together with energy, timing, channel occupancy, and signal quality; the final threshold is fixed by calibration-like control acceptance.

5.4 Multilayer optimization

The design scan compares one-, two-, three-, and four-layer cadmium configurations, together with three-layer gadolinium and stainless-steel variants. The HENSA outdoor neutron field is propagated through the full detector-response chain, including quenching, light attenuation, waveform formation, and veto-peak reconstruction. The scan exporter retains every entry in each processed AnalysisTree; it applies no additional one-track, 27keV, fiducial, or X-ray-topology selection. The denominator is therefore a response-level analysis-entry population, not the conservative reference selection of Section 6. All efficiencies in this section are conditional on that exported population.

The distinction between performance quantities is essential. The probability of any reconstructed veto tag tests whether the geometry produces an observable response. Threshold rejection tests a particular visible-energy operating point. The aggregate classifier adds multiplicity and signal-quality information at a fixed calibration acceptance. None of these conditional quantities is an absolute neutron background level.

QuantityDenominatorNoise / acceptance treatmentPurpose
Any veto tagExported HENSA AnalysisTree entriesNo accidental overlayGeometry response
10MeV threshold rejectionSame response-level sampleReconstructed visible-energy proxy; no accidental overlayComparable layer operating point
Aggregate classifier rejectionThree-layer response sample with calibration-noise overlayThreshold fixed at 90% calibration-like control acceptanceOverlaid design diagnostic
Absolute background levelGenerated exposure and source normalizationSource-appropriate veto credit with prompt/delayed splitReported in Section 6
Table 5.2: Performance definitions used in the veto chapter. Percentages with different denominators or accidental treatments are not interchangeable.
Figure 5.10: HENSA geometry-response scan over all exported processed analysis entries. The left panel shows conditional neutron-event rejection versus reconstructed visible-energy threshold; the shaded interval marks the 10 – 15 MeV -equivalent region motivated by calibration-like control acceptance. The right panel compares the probability of any reconstructed veto tag and the 10 MeV -equivalent rejection. These quantities are not candidate-background efficiencies under the conservative reference selection.
Figure 5.10: HENSA geometry-response scan over all exported processed analysis entries. The left panel shows conditional neutron-event rejection versus reconstructed visible-energy threshold; the shaded interval marks the 1015MeV-equivalent region motivated by calibration-like control acceptance. The right panel compares the probability of any reconstructed veto tag and the 10MeV-equivalent rejection. These quantities are not candidate-background efficiencies under the conservative reference selection.

The any-tag fraction increases from 0.465 for one layer to 0.675, 0.741, and 0.771 for two, three, and four cadmium layers, respectively. At a 10MeV-equivalent threshold, the corresponding conditional rejections are 0.150, 0.292, 0.376, and 0.425. Thus the third layer provides a substantial improvement, whereas the fourth adds 0.030 in any-tag probability and 0.049 in threshold rejection while increasing the channel count from 59 to 79. The gadolinium variant is similar to cadmium for these aggregate observables, whereas stainless steel provides substantially weaker tagging.

Generated-primary normalization gives a processed-analysis-entry probability of approximately 2.2×105 for each of the one- through four-layer cadmium samples. The active layers improve the conditional tagging response; within the statistical precision of this diagnostic, they do not reduce the population reaching the response-level analysis denominator. The exposure audit and full diagnostic projections are given in Appendix Section A.8.

Each additional layer also adds channels that can contribute accidental activity. An illustrative overlay model samples the measured calibration-triggered veto activity and scales the number of independent accidental opportunities with the active layer count. It is used to expose the trade-off, not as a final dead-time measurement.

Figure 5.11: Illustrative noise-aware layer trade-off at a 10 MeV -equivalent threshold. Accidental activity is sampled from calibration-triggered veto data and scaled with layer count. At the nominal overlay scale, the four-layer configuration retains the largest simple relative 𝑆 / 𝐵 proxy but has lower calibration-like control acceptance than the three-layer configuration. The model therefore exposes the performance–acceptance trade-off; it does not statistically select three layers.
Figure 5.11: Illustrative noise-aware layer trade-off at a 10MeV-equivalent threshold. Accidental activity is sampled from calibration-triggered veto data and scaled with layer count. At the nominal overlay scale, the four-layer configuration retains the largest simple relative 𝑆/𝐵 proxy but has lower calibration-like control acceptance than the three-layer configuration. The model therefore exposes the performance–acceptance trade-off; it does not statistically select three layers.

The accidental-overlay model does not, by itself, favor three layers: over the scanned noise range, the four-layer configuration retains the largest simple sensitivity proxy. At the nominal scale, that proxy is 1.274 for four layers and 1.232 for three layers, while the corresponding calibration-like control acceptances are 0.860 and 0.897. Three layers are therefore an engineering and channel-count compromise made after most of the conditional tagging gain has been obtained, not a statistical optimum of this simplified model. The small separation between the three- and four-layer sensitivity proxies should be compared with the response and cascade-model variations before assigning a preferred physics operating point. For an expected background of only a few counts, an expected Poisson-likelihood limit with signal and live-time losses is more appropriate than 𝑆/𝐵.

5.5 Prototype geometry and synchronized readout

5.5.1 Geometry and prototype construction

The 59-panel simulated design is organized into Top, Bottom, Left, Right, Front, and Back groups, with three active stages per group. The nominal symmetric arrangement would contain 60 panels, but one short inner-top panel is absent to accommodate mechanical and service constraints. Thin cadmium sheets are placed between neighboring active stages so that moderated neutrons can capture near a scintillator.

(a) Expanded simulated geometry.
(a) Expanded simulated geometry.
(b) Commissioned surface prototype.
(b) Commissioned surface prototype.
Figure 5.12: Selected three-layer veto geometry and its IAXO-D0 prototype implementation. The left panel shows the 59-panel design-study geometry; the right panel shows the 57-panel surface setup used for the published experimental validation [17].
GroupLayersPanels per layerLengthsDesign note
Top33 / 4 / 480 / 150 cmOne short inner panel is omitted, giving 59 rather than 60 panels.
Bottom34 / 4 / 4150 cmFull long-panel coverage below the detector.
Left33 / 3 / 3150 cmLong side coverage.
Right33 / 3 / 3150 cmLong side coverage.
Front33 / 3 / 380 cmShort modules on the beam-pipe side.
Back33 / 3 / 380 cmShort modules on the rear side.
Table 5.3: Definition of the selected 59-panel design-study geometry. The commissioned IAXO-D0 prototype used 150 cm and 65 cm bars in a 57-signal implementation, whereas the simulated design uses 80 cm short modules.

The prototype reused NE-110 plastic scintillator bars from a former time-of-flight spectrometer [111]. The bars were cut to 150cm and 65cm, with a 20×5cm2 cross-section, and coupled to photomultiplier tubes through light guides. The physical prototype used 1mm cadmium sheets between neighboring panels and layers, realizing the same capture-stage concept as the simulation [17]. Atmospheric muons provided the gain-equalization reference. For a flat 5cm plastic panel, the most probable through-going-muon signal corresponds to approximately 10MeV of visible energy; this anchors the 1015MeV-equivalent region used in the threshold studies [17]. Measurements on the long bars showed that the collected light at the far end can be reduced by approximately a factor of two, so the reconstructed energy is an attenuation- and calibration-dependent proxy rather than a local calorimetric measurement [17].

5.5.2 Electronics and timing

The Micromegas and veto branches each use four AGET application-specific integrated circuits and are synchronized by a common trigger and clock [17, 81]. The circular buffer contains 512 samples, allowing the trigger position and window length to differ between detector branches. For the prototype-like veto configuration, the total acquisition is 100𝜇s with the Micromegas trigger 30𝜇s after the start. The trigger-centered convention used below is therefore approximately [30,+70]𝜇s: near-zero activity provides the prompt tag, while capture-related peaks at positive times contribute to the post-trigger tag only if they occur before the +70𝜇s waveform boundary.

Figure 5.13: Synchronized prototype readout. The Micromegas and veto systems use separate AGET branches with detector-specific timing settings, followed by offline event building in a common trigger-centered coordinate system.
Figure 5.13: Synchronized prototype readout. The Micromegas and veto systems use separate AGET branches with detector-specific timing settings, followed by offline event building in a common trigger-centered coordinate system.

The long veto window is a deliberate part of the neutron-sensitive design. It retains prompt muon and shower activity as well as a substantial fraction of the cadmium-capture tail. The window cannot be widened on capture retention alone, because every additional interval also increases random-coincidence and live-time costs.

5.6 Waveform observables and conditional performance

5.6.1 Detector response and reconstructed observables

The simulation is propagated beyond deposited energy to the quantities used in data. Energy deposits are quenched, attenuated according to their propagation distance, converted into channel signals, shaped, digitized, and processed by the REST peak finder. This common representation permits the same peak time, amplitude, multiplicity, and channel-quality definitions to be used for simulated and experimental events. The common reconstruction-process inventory is given in Table A.1, while veto-specific campaign and response settings are documented in Appendix Section A.6.

Figure 5.14: Signal-processing chain from Geant4 deposits to analysis-level veto observables. Quenching, light attenuation, waveform formation, and peak reconstruction are included before the veto decision.
Figure 5.14: Signal-processing chain from Geant4 deposits to analysis-level veto observables. Quenching, light attenuation, waveform formation, and peak reconstruction are included before the veto decision.

Muon-induced events are predominantly prompt, high-amplitude, and track-like across several panels. HENSA neutron-induced events are more heterogeneous: they contain softer prompt activity, a broader post-trigger tail, and less regular channel correlations. These differences motivate aggregate observables rather than a single opposite-panel coincidence.

Figure 5.15: Waveform-level veto observables for aligned muon and HENSA-neutron simulation samples. Muons concentrate in the prompt region, whereas neutron-induced events populate a broader post-trigger tail and exhibit larger late-window multiplicity. Each distribution is conditional on its stated simulated sample and is not a relative source-rate comparison.
Figure 5.15: Waveform-level veto observables for aligned muon and HENSA-neutron simulation samples. Muons concentrate in the prompt region, whereas neutron-induced events populate a broader post-trigger tail and exhibit larger late-window multiplicity. Each distribution is conditional on its stated simulated sample and is not a relative source-rate comparison.

5.6.2 Calibration-controlled classifier hierarchy

An initial design-facing classifier combined five aggregate reconstructed quantities and established that waveform multiplicity and signal-quality information outperform a single visible-energy threshold. That event-random study is retained in Appendix Section A.4 as a historical benchmark. The run-disjoint study uses the higher-statistics aligned sample of 20,274 HENSA neutron events and 139,454 experimental calibration events. One independently sampled calibration peak train is concatenated with each simulated neutron peak train. This preserves correlations within the measured activity but does not reproduce peak merging, saturation or baseline shifts that can occur when waveforms overlap before peak finding. The control class is the calibration-triggered activity itself, and every score is treated as an ordering variable rather than an event-by-event neutron probability.

The analysis is organized into three implemented levels of increasing opacity. The level-1 boosted decision tree uses six directly interpretable aggregate observables: total reconstructed veto-energy proxy, peak count, unique panel count, active face count, active layer count, and maximum peak amplitude. The level-2 model adds disjoint waveform-window summaries, face and layer occupancies, amplitude fractions and entropies, opposite-face activity, and face–time and layer–time interactions. The level-3 model replaces these engineered spatiotemporal quantities with a two-channel panel–time representation and combines a shared temporal convolution, a relational panel graph, attentive pooling, and the six level-1 aggregates. No Micromegas energy or topology observable enters any level of this classifier hierarchy. Simulation and prototype channel identifiers are first mapped with separate readout maps to a common (face,layer,panel) coordinate; this prevents identical electronic identifiers in the two readouts from being assigned to different physical panels after noise overlay.

The partition is chronological and run-disjoint. Calibration runs 1335, 1339, and 1342 train the models, run 1345 fixes the score thresholds at nominal 90% and 95% control acceptance, and run 1347 is used once for the held-out evaluation. An overlaid neutron event follows the partition of its sampled calibration event, preventing empirical-noise leakage. The mapped waveform coordinate is divided into pre-control bins 0–189, prompt bins 190–210, early post-trigger bins 211–239, capture bins 240–280, and late-tail bins 281–511. The following hierarchy results retain the original response and feature definition as a historical comparison. A subsequent boundary check rejects peaks outside the stored 512-bin record instead of clipping them to its endpoints; the bounded frozen-model comparison is given in Appendix Section A.4.1.1. The physical acquisition support and calibrated total-energy response remain unresolved for this flat input, so the comparison does not establish an absolute neutron-veto efficiency.

ClassifierInputsHeld-out AUCNeutron rejection at nominal 90% control acceptance
Level 1: aggregate60.91183.71.0+0.9%
Level 2: aggregate + time430.91384.01.0+0.9%
Level 2: aggregate + capture delay240.91384.21.0+0.9%
Level 2: aggregate + position490.91284.31.0+0.9%
Level 2: full spatiotemporal1640.91184.21.0+0.9%
Level 3: hybrid neural modeltensor + 60.91384.11.0+0.9%
Full model without total energy1630.89679.11.1+1.1%
Table 5.4: Historical run-disjoint HENSA-neutron rejection for the three-level classifier hierarchy, using the original aligned peak response. Uncertainties are exact two-sided 90% Clopper–Pearson intervals on 4035 held-out neutron events. The final row removes the total-energy proxy as a domain-sensitivity test. The effect of rejecting out-of-record peaks is reported separately in Appendix Section A.4.1.1.
Figure 5.16: Run-disjoint comparison of the aggregate and engineered spatiotemporal veto classifiers. The left panel gives the held-out acceptance–rejection curves; the right panel gives the operating point fixed at nominal 90% calibration-control acceptance on an independent run. Error bars are exact 90% confidence intervals. The dashed curve removes the total reconstructed energy to expose the classifier’s response to timing and position alone.
Figure 5.16: Run-disjoint comparison of the aggregate and engineered spatiotemporal veto classifiers. The left panel gives the held-out acceptance–rejection curves; the right panel gives the operating point fixed at nominal 90% calibration-control acceptance on an independent run. Error bars are exact 90% confidence intervals. The dashed curve removes the total reconstructed energy to expose the classifier’s response to timing and position alone.

The full level-2 model improves the held-out rejection by only 0.47 percentage points relative to level 1, with a central 90% paired-bootstrap interval of 0.07–0.82 percentage points, while its AUC difference is consistent with zero. Moreover, the nominal level-1 calibration acceptance varies from 88.2% to 90.9% across the five runs, whereas the full model varies from 84.4% to 90.5%. The added dimensions therefore amplify one run-dependent control shift without producing a material rejection gain. The total-energy removal test retains 79.1% rejection, demonstrating that timing and position contain independent neutron-sensitive information, but permutation tests show that total visible energy remains the dominant nominal input. Level 1 is consequently retained as the reference classifier for this historical response study; the additional levels test information content without establishing a more robust physical operating point.

5.6.2.1 Capture-delay signature

The small gain from the generic timing vector does not imply that neutron-capture timing is physically uninformative. It answers a narrower question: once total reconstructed veto energy, amplitude, and multiplicity are known, the original broad waveform summaries add little to the global binary classifier. The delayed-capture hypothesis is therefore tested separately, distinguishing the physical capture time from the reconstructed peak time and from an accidental late pulse.

Figure 5.17: Role of capture-delay information in the veto analysis. Geant4 history information is used only to define and interpret the physical timing intervals. The classifiers receive reconstructed VETO observables: the six level-1 aggregates define the nominal selector, while the capture-specific diagnostic combines trigger-relative peak timing with delayed peak amplitude and detector segmentation.
Figure 5.17: Role of capture-delay information in the veto analysis. Geant4 history information is used only to define and interpret the physical timing intervals. The classifiers receive reconstructed VETO observables: the six level-1 aggregates define the nominal selector, while the capture-specific diagnostic combines trigger-relative peak timing with delayed peak amplitude and detector segmentation.

Geant4 truth is used only to establish the expected timescale. For 55,581 positive-delay cadmium captures in the three-layer HENSA sample, the median delay from the first gas deposit is 25.6𝜇s, and a truncated-exponential fit to the 10300𝜇s tail gives 𝜏=46.9𝜇s. The cumulative fractions captured after the trigger and no later than 50, 70, and 100𝜇s are 69.7%, 79.7%, and 88.8%, respectively, using all records with 0<Δ𝑡gas<2000𝜇s as the denominator. The narrower observable interval 10<Δ𝑡gas50𝜇s contains 41.4% of those records; another 28.3% occur within the first 10𝜇s and are not part of the delayed window. This truth information fixes the physical interpretation and the candidate window; it is never supplied to an event classifier.

The historical feature construction maps waveform bin 𝑖 to an aligned time coordinate through

𝑡rel=0.2(𝑖231)𝜇s.

(5.1)

The offset 231 aligns prompt populations empirically; it is not an independent calibration of the hardware-trigger position. In this coordinate, bins 0–511 span [46.2,+56.0]𝜇s, whereas the simulated peak population in the inspected flat neutron input ends at +41.6𝜇s. Consequently, the nominal capture-like interval below is only partly populated by physical simulated peaks; calibration overlay can still populate its remaining bins. The observable prompt interval is 10𝑡rel5𝜇s, while the capture-like interval is 10𝑡rel50𝜇s. The 510mus gap separates the nominal prompt and delayed intervals, but does not by itself bound pulse tails or peak-finding leakage. The explicit timing coordinates are the first delayed-peak time, the interval from the last prompt peak to the first delayed peak, and the delayed amplitude-weighted time mean and width. Binary indicators record whether prompt and delayed peaks exist, so the numerical zero assigned to a missing time coordinate cannot be interpreted as a physical zero delay. The capture-like signature then adds prompt and delayed multiplicities and amplitudes, panel, face, and layer occupancies, delayed fractions, and the number of panels that become active only after the prompt interval.

Figure 5.18: Space-time structure of a representative held-out neutron event with prompt and delayed reconstructed VETO activity. Each marker is a reconstructed peak, its area encodes amplitude, and its vertical coordinate is the common semantic panel used to align simulation and calibration-overlay channels. The delayed response activates panels beyond those present in the prompt interval, illustrating why the signature uses the prompt-to-delayed sequence together with spatial multiplicity rather than treating any late peak as a neutron tag. The event is selected as the closest robust multivariate representative to the median of its predeclared category.
Figure 5.18: Space-time structure of a representative held-out neutron event with prompt and delayed reconstructed VETO activity. Each marker is a reconstructed peak, its area encodes amplitude, and its vertical coordinate is the common semantic panel used to align simulation and calibration-overlay channels. The delayed response activates panels beyond those present in the prompt interval, illustrating why the signature uses the prompt-to-delayed sequence together with spatial multiplicity rather than treating any late peak as a neutron tag. The event is selected as the closest robust multivariate representative to the median of its predeclared category.

A time-scrambled event-mixing test isolates the importance of the sequence. It draws complete veto-peak trains from other simulated neutron events and circularly shifts them within the acquisition interval, preserving energy, multiplicity, and spatial correlations while destroying their timing relation to the Micromegas trigger. This neutron-derived null tests the timing sequence; it does not bound the experimental accidental rate. The resulting trade-off is summarized in Table 5.5.

Observable requirementNeutron acceptanceTime-scrambled acceptanceRatio
At least one 1050𝜇s peak52.4%56.3%0.93
Prompt followed by a 1050𝜇s peak46.4%14.3%3.25
Prompt + delayed activity in at least two veto groups24.6%7.3%3.36
Previous row + delayed peak-energy proxy above 104 analysis units14.9%4.4%3.41
Table 5.5: Observable capture-delay signatures in 879 simulated neutron events. The time-scrambled column is the mean over 250 circularly shifted realizations of neutron-derived peak trains. The ratio divides neutron acceptance by this mean; it is not a measured background rejection or neutron purity. A late peak alone is not neutron-specific; the discriminating information is the trigger-correlated prompt–delayed sequence supplemented by segmentation and the summed delayed VETO peak-energy proxy.
Figure 5.19: Capture-delay validation from physical truth to reconstructed observables. Left: cumulative fraction of simulated cadmium captures from zero delay to the indicated upper boundary, using the 0 < Δ 𝑡 gas < 2000 𝜇 s records as the denominator; the 10 – 50 𝜇 s observable interval is shaded and the prototype + 70 𝜇 s boundary is marked. The cumulative value at 50 𝜇 s is not the acceptance of the shaded interval. Center: neutron acceptance versus acceptance of the neutron-derived time-scrambled control as prompt context, segmentation, and the summed delayed VETO peak-energy proxy are added. Right: held-out rejection at thresholds fixed to nominal 90% calibration-control acceptance on an independent run. Exact 90% intervals are shown for finite neutron samples; horizontal uncertainties in the center panel span the central 90% of the mixed realizations.
Figure 5.19: Capture-delay validation from physical truth to reconstructed observables. Left: cumulative fraction of simulated cadmium captures from zero delay to the indicated upper boundary, using the 0<Δ𝑡gas<2000𝜇s records as the denominator; the 1050𝜇s observable interval is shaded and the prototype +70𝜇s boundary is marked. The cumulative value at 50𝜇s is not the acceptance of the shaded interval. Center: neutron acceptance versus acceptance of the neutron-derived time-scrambled control as prompt context, segmentation, and the summed delayed VETO peak-energy proxy are added. Right: held-out rejection at thresholds fixed to nominal 90% calibration-control acceptance on an independent run. Exact 90% intervals are shown for finite neutron samples; horizontal uncertainties in the center panel span the central 90% of the mixed realizations.

The seven delay coordinates alone reject 49.8% of the held-out neutrons, and the complete 18-variable capture signature rejects 74.1%. Thus, capture timing is informative but is not a universal neutron tag: 24.2% of the selected events containing a truth-level cadmium capture have no reconstructed veto peak (23.8% of all selected events), and late activity also occurs in calibration data. Adding the capture signature to the six level-1 aggregates raises the held-out rejection from 83.69% to 84.19%, corresponding to 20 additional rejected neutrons among 4035. The paired-bootstrap gain is 0.50 percentage points, with a central 90% interval of 0.200.82 percentage points, whereas the AUC difference remains compatible with zero. The minimum control acceptance across the five runs decreases from 88.20% for level 1 to 86.33% for the combined model. The delay is therefore retained as an important, physically interpretable level-2 signature and as a central input to the sequence model, while the aggregate level-1 BDT remains the nominal selector.

5.6.2.2 Level-3 spatiotemporal neural model

The level-3 input has shape 59×512×2: for every common semantic panel and aligned waveform bin, the two channels contain summed peak amplitude and peak occupancy. The simulation-only panel absent from the prototype is masked, and no run identifier, raw channel identifier, overlay identifier, TPC quantity, or Geant4 truth enters the network. A shared temporal convolution encodes each panel; two graph blocks exchange information separately between same-face neighbors, adjacent layers, and opposite faces before attentive pooling. Four predeclared variants distinguish a scalar multilayer perceptron, a tensor-only temporal convolution, a tensor-plus-graph model, and the final hybrid with the six level-1 aggregates.

The tensor-only models reject only 71.2% and 72.6% of the held-out neutrons, respectively, showing that the present sparse peak tensor does not replace the global energy summaries. The hybrid reaches 84.1%, only 0.42 percentage points above level 1; the central paired-bootstrap 90% interval is 0.10–0.74 percentage points, while the AUC difference remains compatible with zero. The frozen score is insensitive at the sub-percentage-point level to shifts of up to five waveform bins and to 10% amplitude rescaling in neutron rejection. Masking one representative panel per face–layer cell changes the rejection by at most 0.37 percentage points, but can reduce control acceptance by 1.25 percentage points. Most importantly, the nominal 90% control acceptance falls to 83.4% in calibration run 1339, below the 84.4% minimum of level 2 and the 88.2% minimum of level 1. The neural model therefore confirms the performance ceiling without satisfying the predeclared run-stability gate.

Figure 5.20: Run-disjoint level-3 comparison at the independently frozen nominal 90% control-acceptance threshold (left) and perturbation envelope of the hybrid neural model (right). The tensor-only temporal and graph branches underperform the aggregate models; the hybrid recovers the aggregate performance but does not materially exceed it. Perturbation intervals span all declared time shifts, amplitude scalings, or representative single-panel masks.
Figure 5.20: Run-disjoint level-3 comparison at the independently frozen nominal 90% control-acceptance threshold (left) and perturbation envelope of the hybrid neural model (right). The tensor-only temporal and graph branches underperform the aggregate models; the hybrid recovers the aggregate performance but does not materially exceed it. Perturbation intervals span all declared time shifts, amplitude scalings, or representative single-panel masks.

The frozen level-1 model is then applied to the exact 249 conservative-reference HENSA candidates from Section 6.6.3. The three causal-history activation events are removed before prompt scoring and receive no veto credit. For each of the remaining 246 candidates, 200 independent calibration events are overlaid. The mean prompt rejection is 91.14%, leaving 21.8 candidates on average; the central 90% range over accidental overlays is 19–25 survivors. The reproducible first overlay realization leaves 22 candidates, corresponding to 𝐵prompt=(3.020.98+1.29)×107countskeV1cm2s1 with a two-sided 90% Garwood interval. The full level-2 model leaves 20.6 candidates on average, a reduction of only 1.2 events relative to level 1. The level-3 hybrid leaves 21.0 candidates on average, with a central 90% overlay range of 18–24, and the first realization leaves 20. This is less selective than level 2 and only 0.8 event below level 1, so it does not alter the controlling prompt-background estimate. The delayed-activation channel remains separate at the one-sided saturation bound reported in Section 6.7. The training, readout maps, run partitions, fixed models, candidate identities, and overlay draws are documented in Appendix Section A.4.

5.7 Experimental validation with IAXO-D0

5.7.1 Published surface cut flow

The commissioned prototype was operated with IAXO-D0 at surface level in Zaragoza for an effective 52.1days [17]. The published selection uses a 27keV region of interest and a 9mm-radius focal region. This should not be confused with the 10mm fiducial radius used by the current conservative background-analysis contract.

StageSelectionEventsBackground level
Raw acquisitionFull detector and energy range1,305,996
Micromegas selection27keV, 9mm focal region, X-ray-like topology257
Prompt vetoPrompt muon-like veto discrimination569.782.05+2.44×107
Advanced vetoMultiplicity and activity in prompt and post-trigger sub-windows498.561.91+2.30×107
Table 5.6: Published IAXO-D0 surface cut flow [17]. Background levels are in countskeV1cm2s1. The central values reproduce the publication; the asymmetric uncertainties shown here are exact 90% Poisson confidence intervals derived from the reported event counts. Calibration efficiency changes from 81.9% after the Micromegas selection to 79.4% after the full veto selection.

The prompt selection removes 201/257=78.2% of the Micromegas-selected events and is the dominant surface-background rejection. The advanced selection removes 7/56=12.5% of the post-prompt sample; the exact two-sided 90% binomial interval is 6.0%22.2%. The calibration efficiency after all cuts is 79.4/81.9=97.0% relative to the Micromegas-selection reference. No separate prompt-stage calibration efficiency was reported, so this ratio is not described as retention relative to the prompt veto.

Figure 5.21 replots the reported cut-flow values in a terminology consistent with the present analysis. The stage historically called a “neutron cut” is labeled a neutron-sensitive control selection because the data do not establish the origin of individual rejected events.

Figure 5.21: IAXO-D0 surface cut-flow summary reconstructed from the published event counts and background levels in Ref. [ 17 ]. Panel (a) shows the cumulative candidate population after the Micromegas, prompt-veto, and advanced veto selections. Panel (b) gives the corresponding exposure-normalized background levels after the two veto stages, with exact 90% Poisson intervals. The final stage is described as a neutron-sensitive control selection; it does not provide event-by-event neutron identification.
Figure 5.21: IAXO-D0 surface cut-flow summary reconstructed from the published event counts and background levels in Ref. [17]. Panel (a) shows the cumulative candidate population after the Micromegas, prompt-veto, and advanced veto selections. Panel (b) gives the corresponding exposure-normalized background levels after the two veto stages, with exact 90% Poisson intervals. The final stage is described as a neutron-sensitive control selection; it does not provide event-by-event neutron identification.

The data support the hierarchy expected from the design. Prompt activity provides the large muon-like rejection, and the long waveform supplies an additional handle on non-prompt or multiplicity-rich activity with a small calibration penalty. The seven additionally rejected events do not measure a neutron efficiency because the post-prompt sample contains an unknown mixture of neutron-induced, proton-induced, environmental, activation, and instrumental backgrounds.

5.7.2 Late-window control population

An exploratory thesis extension ranks the feature-complete flattened prototype sample against calibration accidentals, muon simulation, and HENSA-neutron simulation. After removing prompt-muon-like and burst-like events, a frozen score orders the remaining events by their similarity to the neutron-plus-noise template relative to calibration and muon controls. Its threshold is fixed at the upper 1% tail of the clean calibration distribution. The score is deliberately used as a control-population selector rather than as an event-by-event neutron probability; its exact construction, preprocessing, and time-window definitions are documented in Appendix Section A.5.

Figure 5.22: Late-window score distributions and the 99th-percentile calibration threshold, which retains the upper 1 % calibration tail. A large score ranks an event closer to the neutron-plus-noise template but does not establish its physical origin.
Figure 5.22: Late-window score distributions and the 99th-percentile calibration threshold, which retains the upper 1% calibration tail. A large score ranks an event closer to the neutron-plus-noise template but does not establish its physical origin.
SampleEventsPrompt tagClean after promptMedian scoreSelected
Calibration13945411.00%1234413.801235
Flat experimental background77775291.13%659902.103877
Muon+noise simulation2156997.98%4332.4712
Neutron+noise simulation2027451.95%96972.701722
Table 5.7: Population summary for the frozen late-window score. The selected column applies the 99th-percentile calibration threshold after prompt and burst removal. The entries are selection counts within the separate stated denominators, not inferred particle fractions.

The score selects 3877 experimental events, corresponding to 0.50% of the full flattened input and 5.88% of its prompt-suppressed subset. The selected data differ from the calibration control and have passed prompt-muon-like rejection, but they are not quantitatively described by the present HENSA template. In particular, their veto multiplicity, late-window energy, and Micromegas topology differ substantially from the selected neutron-plus-noise simulation. The defensible result is therefore the observation of a calibration-controlled late-window population, not a neutron fraction or a background-model normalization. The detailed template-mismatch table and channel correlations remain in Appendix Section A.5; the event rasters and Micromegas projections are collected in Appendix Section A.10.

5.7.3 Limits of experimental neutron identification

The frozen score identifies an experimentally distinct population enriched in neutron-sensitive observables, but it does not identify a sample of neutrons with known purity. At the threshold accepting 1% of the clean calibration control, 5.88% of the prompt-suppressed background population is selected. This establishes a difference between the score distributions of background and calibration events. The score also contains Micromegas energy, and its overlapping time windows and run-dependent control populations can produce a difference without a unique new particle component. The comparison therefore does not exclude ordinary accidental or instrumental explanations and does not establish individual event origins.

The principal obstacle is the quantitative mismatch with the available neutron template. The selected data contain a median of 47 reconstructed veto peaks, compared with 9 in the selected neutron-plus-noise simulation, and their median late-window energy proxy is approximately 2.5 times larger. Their Micromegas energy, hit multiplicity, and topology also differ substantially. Consequently, interpreting all selected events as neutrons would contradict the measured observables, while converting the selected count into a neutron rate would import an unsupported template-purity assumption.

Several effects prevent a stronger attribution in the present study. No neutron-source or otherwise tagged neutron data are available for the commissioned configuration. The HENSA production represents an outdoor spectral model rather than a run-matched experimental calibration, and the 59-panel simulation does not reproduce every gap, threshold, gain, unstable channel, and time-dependent condition of the 57-panel prototype. The experimental sample can also contain residual proton- and photon-induced showers, activation products, correlated electronics activity, and other instrumental populations for which complete normalized templates are not available. Finally, a late peak is not neutron-specific: accidental activity is common, and long-lived activation is causally distinct from an in-waveform capture response.

The existing data can strengthen the control result through run- and channel-state matching, matching in Micromegas energy, a veto-only score, and time-shifted or off-time control windows. Score construction and threshold setting should use different runs from the comparison sample. A quantitative neutron fraction would additionally require credible competing templates and independent particle-response constraints, for which tagged neutron data would be particularly useful. Until these conditions are met, the thesis reports a calibration-controlled, neutron-sensitive late-window population rather than an observed neutron count.

5.8 Systematic limitations

Uncertainty classConsequence for the veto conclusion
Surface neutron fieldHENSA constrains the Zaragoza outdoor field; transfer to the planned DESY installation remains site dependent while the detector and site contracts are not frozen.
Hadronic and capture modelsLead secondary production and cadmium gamma partition alter threshold-dependent response. The tested replay gives a model spread; its coverage of thermal transport, the frozen classifier and GeV-scale layer ordering is incomplete.
Visible-energy responseRecoil identity and true step length are not fully represented by the historical quenching approximation. Quenching, effective attenuation, gain and peak finding must be propagated to panel thresholds and selected candidates.
Geometry and channel stateThe 59-panel simulation and 57-panel prototype differ in coverage, channel mapping, gaps, thresholds, and unstable-channel history.
Acquisition and accidentalsPeak-list overlays preserve measured control patterns but omit waveform overlap. Trigger mapping, physical record support, channel rates and online logic determine acceptance loss.
Experimental particle attributionThe late-window selection is enriched relative to calibration controls, but no tagged neutron sample or complete set of normalized competing templates is available. It therefore cannot determine neutron purity or rate.
Delayed activationLong-lived radioactive descendants are not coincident with the initiating shower and cannot receive prompt-veto rejection credit.
Table 5.8: Systematic limitations relevant to the design and prototype validation. Absolute source-normalization uncertainties are propagated in Section 6.

The background model also evaluates cosmic-induced random veto activity in a fixed simulated window, as described in Section 6.6.5. That result is distinct from a measured full-system dead-time estimate, which remains configuration dependent. The present chapter therefore reports calibration acceptance for offline selections and does not infer an unmeasured online live-time correction.

Detector inclination during solar tracking was tested separately for muons and HENSA neutrons. Across the simulated 25 to +25 range, the processed TPC-response acceptance varies at the few-percent level without a resolved monotonic trend. The scan in Appendix Section A.3 is conditional on its historical angular input; it does not close the final tracking-exposure dependence before that input is reconciled.

5.9 Summary and outlook

The shielding studies establish a coherent surface-detector strategy. Lead is indispensable for photon attenuation but does not remove the fast-neutron component; instead, it helps create the correlated secondary shower that an active system can tag. Passive borated-HDPE moderation improves selected configurations only modestly and cannot replace the veto.

The historical prototype design uses three scintillator–cadmium stages around the lead shield. The conditional HENSA response increases strongly through the second and third stages and only modestly in the fourth, while additional channels increase mechanical and accidental-veto costs. Generated-primary normalization shows no resolved reduction of the processed analysis-entry population across the one- through four-layer scan; the observed gain is additional reconstructed tagging information.

The IAXO-D0 data validate the two-level analysis logic. Prompt veto activity removes the dominant muon-like population, and late-window or multiplicity-rich waveform observables provide a smaller additional rejection while retaining 97.0% of the Micromegas-selected calibration efficiency. The simulations establish neutron-sensitive response, whereas the prototype data demonstrate an additional non-prompt rejection handle without identifying the seven rejected events as neutrons. The extended data study similarly isolates a calibration-controlled population enriched in delayed and multiplicity-rich activity, but its mismatch with the HENSA template prevents an experimental neutron count or fraction from being inferred. The present experimental result is additional rejection with unresolved particle attribution. Absolute prompt and delayed background levels are evaluated by the source-normalized analysis in Section 6.

Within the historical simulation-based hierarchy, the six-observable aggregate boosted tree provides the reference result. The engineered timing and segmentation model confirms that these observables contain independent neutron-sensitive information, but its nominal rejection gain is only about half a percentage point and it is less stable across one calibration run. The completed neural study reaches the same conclusion independently: explicit temporal convolution and spatial message passing recover the aggregate performance only when the aggregate branch is restored, while the resulting gain is below one candidate and the run dependence worsens. Greater model complexity is therefore retained as a validation study rather than adopted as the nominal selector. The next useful simulations should first close recoil visibility, acquisition timing and source angular support, then propagate cascade and response alternatives through a frozen selector. A March 2026 collaboration design study explores compact three-layer modules with wavelength-shifting fibers and SiPM readout, including boron-loaded material and cadmium options [112]. That geometry should be compared with the historical PMT implementation under a common mechanical envelope and response model; single-bar test efficiency is not a measurement of complete-veto efficiency.