Binary neutron stars
Dingo analyzes binary neutron star (BNS) events with the DINGO-BNS method of [1]. BNS signals pose two problems for plain NPE. Their long inspirals require a fine frequency resolution, so the data are far larger than for binary black holes, and their chirp mass is measured so precisely that a network covering the full training prior would spend nearly all of its capacity on parameter values excluded by any individual event. DINGO-BNS addresses both with a single device: a chirp-mass proxy that simplifies the data (phase heterodyning) and narrows the effective prior (prior conditioning). The proxy is set per event, so sampling is one pass through the network, with the density preserved and importance sampling available directly.
Phase heterodyning
At BNS frequency resolutions the strain oscillates rapidly in frequency, which makes poor input for a network. Multiplying the data by \(\exp(i \phi(f; \tilde{\mathcal{M}}))\) with the leading-order chirp phase
removes the dominant phase evolution for a reference chirp mass \(\tilde{\mathcal{M}}\) close to the true value. The residual oscillations are slow, and the multibanded frequency domain can then decimate the data far more aggressively.
The reference value \(\tilde{\mathcal{M}}\) is the chirp-mass proxy,
chirp_mass_proxy. Following [1], the network is trained on a
family of restricted chirp-mass priors centered on \(\tilde{\mathcal{M}}\) and
conditioned on \(\tilde{\mathcal{M}}\) (prior conditioning). In
training, \(\tilde{\mathcal{M}} = \mathcal{M} + \epsilon\) with \(\epsilon\) drawn from a
narrow kernel, so the kernel’s support is the restricted prior. At inference the proxy
is set per event, from a search trigger or the
chirp-mass scan, and one network pass gives samples with their
density. The implementation reuses the proxy machinery of GNPE, without
its Gibbs iteration. A prior-conditioned model records the kernel and the phase order
in its metadata under chirp_prior_conditioning, and the
sampler context reads this to prepare data
as a function of chirp_mass_proxy. Heterodyning is applied to the raw strain before
decimation (the two operations do not commute).
Prior conditioning
The proxy plays a second role. Because the network conditions on
\(\tilde{\mathcal{M}}\), it effectively learns a family of posteriors under narrow
chirp-mass priors centered on the proxy, \(q(\theta | d, \tilde{\mathcal{M}})\).
Setting the proxy at inference selects the member of the family appropriate to the
event, so one network amortizes over events while retaining the resolution of an
event-specific narrow prior. The network infers the offset
delta_chirp_mass \(= \mathcal{M} - \tilde{\mathcal{M}}\) rather than the chirp mass
itself; the chain reconstructs the physical value with a ProxyOffsetReparam step.
In the notation of [1], the reference chirp mass
\(\tilde{\mathcal{M}}\) is chirp_mass_proxy, the prior half-width
\(\Delta\mathcal{M}\) is the half-width of the kernel, the offset
\(\delta\mathcal{M}\) is delta_chirp_mass, and the hyperprior over
\(\tilde{\mathcal{M}}\) is the dataset’s chirp-mass prior convolved with the kernel.
The chain:
flowchart TB
pins["DeltaFactor<br/>chirp_mass_proxy"]
flow["FlowFactor<br/>draws delta_chirp_mass, #8230;"]
off["ProxyOffsetReparam<br/>chirp_mass = delta_chirp_mass + chirp_mass_proxy"]
out(["samples + log_prob"])
ctx["GWSamplerContext<br/>heterodyne #8594; decimate #8594; whiten"]
pins --> flow --> off --> out
pins -. "chirp_mass_proxy" .-> ctx
ctx -. "prepared data" .-> flow
classDef step fill:#dbe9f6,stroke:#2980b9,color:#1a1a1a
classDef reparam fill:#e2f0e6,stroke:#27ae60,color:#1a1a1a
classDef ctxstyle fill:#f4f4f4,stroke:#8c8c8c,color:#1a1a1a
class pins,flow step
class off reparam
class ctx ctxstyle
The pinned values have a single owner, the chain root, and are recorded with the
samples. The heterodyne receives the proxy through the row-aligned prepared_data
contract of the sampler context. Since the
chain contains no Gibbs block, the samples carry their log probability and
importance sampling proceeds without a density-recovery step.
A network may condition on further context parameters, such as the sky position, and any of them can be pinned in the same way; the frame handling of a pinned right ascension is described under sampling chains.
The chirp-mass scan
When no external estimate of the chirp mass is available, the trigger value can be determined from the data (see the Methods of [1]). The scan sweeps the proxy over the training chirp-mass prior on a grid whose spacing is set by the kernel width, draws a few posterior samples at each grid point in batched network passes over blocks of grid points, evaluates a phase-marginalized likelihood for every draw within the prior (the scan therefore requires a phase-marginalized network), and takes the chirp mass of the maximum-likelihood draw as the trigger value.
The sweep runs on the ordinary chain machinery: a fixed table with one row per grid point roots the chain, the data preparation heterodynes each row at its own proxy value, and the network draws per row. The scan result (trigger value, signal-to-noise ratio, maximum log likelihood, and the scan settings) is recorded in the sampler provenance. For a GW170817-like event the scan costs about a minute of CPU time.
Multibanding heterodyned data
The Dingo-BNS method combines both heterodyning and multi-banding. Heterodyning factors out the dominant frequency evolution, leaving only slow residual oscillations in the waveform. Multi-banding can then be applied to aggressively coarsen sampling at higher frequencies. Note that since decimation does not commute with heterodyning, decimation nodes must be chosen after heterodyning.
The multi-banding nodes can be specified manually in the waveform dataset config file, or this can be done automatically using the
band tool. The CLI tool dingo_generate_multibanded_domain starts from a uniform frequency domain and decimates
test waveforms until their mismatch with the originals meets a target. When the dataset settings contain phase_heterodyning, the tool heterodynes the waveforms first.
There is one more subtlety. The network never sees data heterodyned at the true chirp mass, only at the proxy, which can sit anywhere within the kernel width of the true value (\(\pm 0.005\,M_\odot\) in the example). This matters because the residual oscillation depends not just on how far off the proxy is, but on which side it lies: the offset adds a term to the residual phase that flips sign with it, and this either adds to the post-Newtonian remainder or partially cancels it. One side of the kernel therefore leaves a faster-oscillating residual than the other, and which side is worse can vary with frequency. To make sure the bands work for both, pass --chirp_mass_proxy_offset set to the kernel half-width. The tool then heterodynes alternate test waveforms at the chirp mass plus and minus this offset, and enforces the mismatch target on whichever side is worse. The example shows the full command.
Tidal approximants
Tidal approximants such as IMRPhenomXP_NRTidalv3 run through the standard LAL
WaveformGenerator in both the uniform and the multibanded frequency domain: the
tidal deformabilities lambda_1 and lambda_2 are inserted into the LAL parameter
dictionary whenever present, and they are ordinary inference parameters, listed with
the others (see the example). The network is phase marginalized, so
the phase is reconstructed synthetically before importance sampling. For these models,
train with spin_conversion_phase: null (Bilby’s convention): a phase shift is then a
global factor \(e^{2i\phi_c}\), so the synthetic phase with approximation_22_mode: true
is exact, and importance sampling reuses its log likelihood. The same holds for a
network that also marginalizes \(\psi\): both angles are then recovered from one waveform
evaluation per sample, with the phase drawn exactly. For a network trained with
spin_conversion_phase: 0.0 the default is the exact mode sum, which for these models
needs a mode_list in the waveform generator settings; approximation_22_mode: true
gives an approximate proposal instead, and importance sampling stays unbiased, at a
lower efficiency. See the synthetic phase section.