Training
Training a network can require a significant amount of time (for production models, typically a week with a fast GPU). We therefore expect that this will almost always be done non-interactively using a command-line script. Dingo offers two options, dingo_train and dingo_train_condor, depending on whether your GPU is local or cluster-based.
Both of these scripts take as main argument a settings file, which specifies options relating to Data pre-processing, training strategy, Neural network architecture, hardware, and checkpointing. They produce a trained model in PyTorch .pt format, and they save checkpoints and the training history. The settings file is furthermore saved within the model files for reproducibility and to be able to resume training from a checkpoint. Finally, all precursor settings files (for the waveform or noise datasets) are also saved with the model.
Settings file
train_settings.yaml file. This is also available in the examples/ folder. The specific settings listed will train a production-size network, taking about a week on an NVIDIA A100. Consider reducing some model hyperparameters for experimentation.data:
waveform_dataset_path: /path/to/waveform_dataset.hdf5 # Contains intrinsic waveforms
train_fraction: 0.95
window:
type: tukey
f_s: 4096
T: 8.0
roll_off: 0.4
domain_update:
f_min: 20.0
f_max: 1024.0
svd_size_update: 200
detectors:
- H1
- L1
extrinsic_prior:
dec: default
ra: default
geocent_time: bilby.core.prior.Uniform(minimum=-0.10, maximum=0.10, name='geocent_time')
psi: default
luminosity_distance: bilby.core.prior.Uniform(minimum=100.0, maximum=1000.0, name='luminosity_distance')
ref_time: 1126259462.391
gnpe_time_shifts:
kernel: bilby.core.prior.Uniform(minimum=-0.001, maximum=0.001)
exact_equiv: True
# chirp_prior_conditioning: # Prior conditioning on the chirp mass (DINGO-BNS); see the BNS page.
# kernel:
# chirp_mass: bilby.core.prior.Uniform(minimum=-0.005, maximum=0.005)
# context_parameters: [ra, dec] # Further conditioning parameters, pinned per event at inference.
inference_parameters: default
model:
posterior_model_type: normalizing_flow
posterior_kwargs:
num_flow_steps: 30
base_transform_kwargs:
hidden_dim: 512
num_transform_blocks: 5
activation: elu
dropout_probability: 0.0
norm: BatchNorm
num_bins: 8
base_transform_type: rq-coupling
embedding_kwargs:
output_dim: 128
hidden_dims: [1024, 1024, 1024, 1024, 1024, 1024,
512, 512, 512, 512, 512, 512,
256, 256, 256, 256, 256, 256,
128, 128, 128, 128, 128, 128]
activation: elu
dropout: 0.0
norm: BatchNorm
svd:
num_training_samples: 20000
num_validation_samples: 5000
size: 200
# Training is divided in stages. They each require all settings as indicated below.
training:
stage_0:
epochs: 300
asd_dataset_path: /path/to/asds_fiducial.hdf5
freeze_rb_layer: True
optimizer:
type: adam
lr: 0.0001
# fused: false # True selects the fused CUDA kernel; see the multi-GPU guide.
scheduler:
type: cosine
T_max: 300
batch_size: 512 # Total effective batch size across all GPUs.
# gradient_updates_per_optimizer_step: 1
# automatic_mixed_precision: False
stage_1:
epochs: 150
asd_dataset_path: /path/to/asds.hdf5
freeze_rb_layer: False
optimizer:
type: adam
lr: 0.00001
scheduler:
type: cosine
T_max: 150
batch_size: 512
# gradient_updates_per_optimizer_step: 1
# automatic_mixed_precision: False
# Local settings that have no impact on the final trained network.
local:
device: cuda # Change this to 'cpu' for training without a GPU.
num_workers: 6
# num_gpus: 1 # Set to >1 to enable multi-GPU (DDP) training. Requires
# freeze_rb_layer: False in all stages. When using
# dingo_train_condor, request_gpus is set automatically.
# ddp_port: 12355 # Rendezvous port for DDP; change when running several jobs on one node.
# torch_compile: False # Fuse the network kernels with torch.compile; see the multi-GPU guide.
# float32_matmul_precision: highest # 'high' enables TensorFloat-32 matmuls (A100+); see the multi-GPU guide.
# wandb:
# project: dingo
# group: my_project
runtime_limits:
max_time_per_run: 36000
max_epochs_per_run: 500
checkpoint_epochs: 10
leave_waveforms_on_disk: True
local_cache_path: tmp
# condor:
# num_cpus: 16
# memory_cpus: 128000
# memory_gpus: 8000
# request_disk: 50GB
The train settings file is grouped into four sections:
data_settings
These settings point to a saved dataset of waveform polarizations and describe the transforms to obtain detector waveforms. A detailed description of these settings is available here.
model
This describes the model architecture, including network type and hyperparameters. All of these settings are described in the section on Neural network architecture.
training
This describes the training strategy. Training is divided into stages, each of which can differ to some extent. Stages are numbered (stage_0, stage_1, …) and executed in this order. Each stage is defined by the following settings:
- epochs
Total number of training epochs for the stage. The network sees the entire training set once per epoch.
- asd_dataset_path
Points to an
ASDDatasetfile. Each stage can have its own ASD dataset, which is useful for implementing a pre-training stage with fixed ASD and a fine-tuning stage with variable ASD.- freeze_rb_layer
Whether to freeze the first layer of the embedding network. This layer is seeded with reduced (SVD) basis vectors, so freezing this layer during pre-training simply projects data onto the basis coefficients. In the fine-tuning stage, when other weights are more stable, unfreezing this can be useful.
- optimizer
Specify optimizer type and parameters such as initial learning rate.
- scheduler
Use a learning rate scheduler to reduce the learning rate over time. This can improve overall optimization.
- batch_size
Total number of training samples per optimizer step. For a training dataset of size \(N\), each epoch consists of \(N /\)
batch_sizeoptimizer steps. When using multiple GPUs, this is the effective batch size across all GPUs; each GPU processesbatch_size/num_gpussamples per step.- gradient_updates_per_optimizer_step
(Optional, default 1) Number of forward–backward passes to accumulate before calling the optimizer. Setting this to \(k\) simulates an effective batch size of \(k \times\)
batch_sizewithout increasing GPU memory usage. This is useful when the desired batch size does not fit in GPU memory.- automatic_mixed_precision
(Optional, default
False) Enable automatic mixed precision (AMP) training. With AMP, the forward pass runs in FP16, while optimizer state and parameter updates remain in FP32. This can roughly halve GPU memory usage and increase throughput on GPUs with Tensor Core hardware (e.g. NVIDIA A100, V100). Requires PyTorch >= 2.0 and a CUDA device.
Important
The stage-training framework allows for separate pre-training and fine-tuning stages. We found that having a pre-training stage where we freeze certain network weights and fix the noise ASD improves overall training results.
local
The local settings are the only group that have no impact on the final trained network. Indeed, they are not even saved in the .pt files; rather they are split off and saved in a new file local_settings.yaml.
- device
cpuorcuda. Training on a GPU with CUDA is highly recommended.- num_workers
Number of CPU worker processes to use for pre-processing training data before copying to the GPU. Data pre-processing (inluding decompression, projection to detectors, and noise generation) is quite expensive, so using 16 or 32 processes is recommended, otherwise this can become a bottleneck. We recommend monitoring the GPU utilization percentage as well as time spent on pre-processing (output during training) to fine-tune this number. When training on multiple GPUs, this is the total number of workers; it is divided equally across the GPUs.
- num_gpus
(Optional, default 1) Number of GPUs to use for training. Setting this to more than 1 enables data-parallel multi-GPU training (PyTorch DDP). The
batch_sizespecified in each training stage is the total effective batch size; it is divided equally across GPUs. When usingdingo_train_condor, the HTCondorrequest_gpusdirective is set automatically from this value. See Multi-GPU training for details.- ddp_port
(Optional, default 12355) Port used for the rendezvous of the DDP processes. When running several multi-GPU jobs on the same node, choose a different port for each to avoid collisions.
- wandb
Settings for Weights & Biases. If you have an account, you can use this to track your training progress and compare different runs.
- runtime_limits
Maximum time (in seconds) or maximum number of epochs per run. Using this could make sense in a cluster environment.
- checkpoint_epochs
Dingo saves a temporary checkpoint in
model_latest.pyafter every epoch, but this is later overwritten by the next checkpoint. This setting saves a permanent checkpoint after the specified number of epochs. Having these checkpoints can help in recovering from training failures that do not result in program termination.- leave_waveforms_on_disk
To improve memory efficiency during training, the waveforms are not loaded into memory at the beginning of training, but separately for each batch during training. When training on a cluster, it is highly recommended to include a local path where the dataset is cached at the beginning of training (see
local_cache_path). If RAM is not an issue, the defaultleave_waveforms_on_disk=Truecan be set toFalse.- local_cache_path
When training on a cluster and loading waveforms during training (i.e.,
leave_waveforms_on_disk=True), the waveform dataset should be copied to the disk storage of the local node at the beginning of training. This prevents unexpected long data loading times during training due to network traffic. Usually, paths for local storage aretmpordev/shm. When submitting the job withcondor,request_disk: 50GBshould be included in thecondorsettings with the requested disk space larger than the size of the waveform dataset used for training.- condor
Settings for HTCondor. The condor script will (re)submit itself according to these options. Available keys are
bid,num_cpus,memory_cpus,memory_gpus,request_disk,requirements, andextra_submit_lines(a list of additional lines written verbatim to the submission file, e.g., for cluster-specific templates). The number of requested GPUs (request_gpus) is derived automatically fromnum_gpusabove and does not need to be specified here.
Multi-GPU training
Dingo supports data-parallel training across multiple GPUs using PyTorch DDP (DistributedDataParallel). Each GPU processes a different shard of the mini-batch simultaneously, effectively scaling throughput with the number of GPUs.
For a detailed practical guide — including how to scale batch_size and learning
rate for actual speedup, how to interpret the training log, and troubleshooting
tips — see Multi-GPU Training.
Enabling multi-GPU training
Set num_gpus in the local section of your settings file:
local:
device: cuda
num_gpus: 4
num_workers: 32 # Total across all GPUs, e.g. 4–8 per GPU
...
The batch_size in each training stage is always the total effective batch size
across all GPUs. Dingo divides it equally, so batch_size must be divisible by
num_gpus. Each GPU therefore processes batch_size / num_gpus samples per step,
which keeps the gradient statistics and learning dynamics identical to single-GPU
training with the same batch_size.
Important
batch_size is the total effective batch size across all GPUs, not the per-GPU
batch size. Increasing num_gpus does not change the effective batch size or
require adjusting any other hyperparameters.
`freeze_rb_layer = True` is currently not allowed since additional changes
would be required to exclude the frozen RB layer from the DDP model wrapper.
Gradient accumulation
For very large effective batch sizes that do not fit in GPU memory, use gradient accumulation alongside multi-GPU training:
training:
stage_0:
batch_size: 4096
gradient_updates_per_optimizer_step: 4 # effective batch = 4096 * 4 = 16384
Each optimizer step accumulates gradients over gradient_updates_per_optimizer_step
forward–backward passes. Multi-GPU and gradient accumulation can be combined freely.
Automatic mixed precision
Enable AMP to reduce GPU memory usage and increase throughput on modern GPUs:
training:
stage_0:
batch_size: 4096
automatic_mixed_precision: True
AMP runs the forward pass in FP16 and the optimizer step in FP32, typically halving memory usage with negligible impact on model quality. It requires a CUDA device and PyTorch >= 2.0.
Using multiple GPUs with dingo_train_condor
When submitting via HTCondor, the request_gpus directive is set automatically from
local.num_gpus. No changes to the condor: block are needed:
local:
device: cuda
num_gpus: 4
condor:
num_cpus: 64
memory_cpus: 256000
memory_gpus: 24000 # Memory per GPU in MB — used for node selection only.
request_disk: 50GB
Note
On clusters with full-node GPU allocations (e.g., MPI-IS), the required HTCondor
template can be passed through with extra_submit_lines, e.g.,
condor:
extra_submit_lines:
- "use template : FullNode"
### Requirements
- PyTorch >= 2.0 (required for `torch.amp`)
- An NCCL-capable build of PyTorch (standard for CUDA-enabled installations)
- One CUDA GPU per process
## Command-line scripts
### `dingo_train`
On a local machine, simply pass the settings file (or checkpoint) and an output directory to `dingo_train`. It will train until complete, or until a runtime limit is reached.
```text
usage: dingo_train [-h] [--settings_file SETTINGS_FILE] --train_dir TRAIN_DIR [--checkpoint CHECKPOINT]
Train a neural network for gravitational-wave single-event inference.
This program can be called in one of two ways:
a) with a settings file. This will create a new network based on the
contents of the settings file.
b) with a checkpoint file. This will resume training from the checkpoint.
optional arguments:
-h, --help show this help message and exit
--settings_file SETTINGS_FILE
YAML file containing training settings.
--train_dir TRAIN_DIR
Directory for Dingo training output.
--checkpoint CHECKPOINT
Checkpoint file from which to resume training.
dingo_train_condor
On a cluster using HTCondor, use dingo_train_condor. This calls itself recursively as follows:
The first time you call it, use the flag
--start-submission. This creates a condor submission filesubmission_file.subthat again calls the executabledingo_train_condor(now without the flag) and submits it. This will rundingo_train_condordirectly on the cluster node that is assigned.On the cluster node,
dingo_train_condorfirst trains the network until done or a runtime limit is reached (be careful to set this shorter than the condor time limit). Then it creates a new submission file that once again callsdingo_train_condor, and submits it. This will resume the run on a new node, and repeat.
usage: dingo_train_condor [-h] --train_dir TRAIN_DIR [--checkpoint CHECKPOINT] [--start_submission]
optional arguments:
-h, --help show this help message and exit
--train_dir TRAIN_DIR
Directory for Dingo training output.
--checkpoint CHECKPOINT
--start_submission
Output
Output from training is stored in the TRAIN_DIR folder passed to the training scripts. This consists of the following:
model_latest.ptcheckpoints every epoch (overwritten);model_XXX.ptcheckpoints whereXXXis the epoch number, everycheckpoint_epochsepochs;model_stage_X.ptat the end of training stageX;history.txtwith columns (epoch number, train loss, test loss, learning rate);svd_L1.hdf5, …, storing SVD basis information used for seeding the embedding network;local_settings.yamlwith local settings for the run (not stored with checkpoints).
The .pt and .hdf5 files may all be inspected using dingo_ls. This prints all the settings, as well as diagnostic information for SVD bases. The saved settings include all the settings provided in the settings file, as well as several derived quantities, such as parameter standardizations, additional context parameters (for GNPE), etc.
Modifying a checkpoint
Occasionally it may be necessary to change a setting of a partially trained model. For example, a model may have been successfully pre-trained, but the fine-tuning failed, and one may wish to change the fine-tuning settings without starting from scratch. Since the model setting are all stored with the checkpoint, they just need to be changed.
The script dingo_append_training_stage allows for appending a model stage or replacing an existing planned stage. It will fail if the stage has already begun training, so be sure to use it on a sufficiently early checkpoint.
usage: dingo_append_training_stage [-h] --checkpoint CHECKPOINT --stage_settings_file STAGE_SETTINGS_FILE --out_file OUT_FILE [--replace REPLACE]
optional arguments:
-h, --help show this help message and exit
--checkpoint CHECKPOINT
--stage_settings_file STAGE_SETTINGS_FILE
--out_file OUT_FILE
--replace REPLACE
For more detailed adjustments to the training settings the script one can use the script compatibility/update_model_metadata.py.
usage: update_model_metadata.py [-h] --checkpoint CHECKPOINT --key KEY [KEY ...] --value VALUE
optional arguments:
-h, --help show this help message and exit
--checkpoint CHECKPOINT
--key KEY [KEY ...]
--value VALUE
Warning
Modifications to model metadata can easily break things. Do not use this unless completely sure what you are doing!