Status¶
gatle-ignite is 0.4.1. This page says what has been run and what has not, because "it should
work" is a weaker claim than "it was run".
Verified¶
Checked on CPU, on every commit.
- Training end to end: the synthetic example drives loss 2.31 → 0.59 and accuracy 0.22 → 0.93.
- Checkpointing:
latestevery epoch,valid_best_resulton each new best, and best-selection by score whenn_saved > 1. - Stopping a diverged run: a NaN loss mid-epoch or a NaN
score_namestops the run and leaves nobestorlatestfrom that epoch. That one rank's NaN stops every rank was run once, on a 2-rank gloo run, and is not in the suite. - Resume restarts at the saved epoch, not from scratch.
- Inference mode loads a checkpoint and evaluates without training, and writes no checkpoints.
- Metric dispatch, composite-loss weighting and iteration gating, and
get_valuepath resolution. - LR schedule shapes:
warmup_epochs = Ngives an N-epoch warmup, and the otherlr_scheduler_paramsreach the scheduler. - Config validation: required fields, field values, overrides.
- Optional deps stay optional: a text-only run imports neither
wandbnorrequests. - Distributed: two-process
glooruns shard the train loader, all-reduce metrics, and on an eval set that does not divide evenly across ranks report the single-process answer. Eval splits are sharded exactly (ExactDistributedSampler), so no sample is scored twice. - Packaging: the wheel, installed into a clean environment, trains a project from a directory
outside the repo with nothing else on
sys.path. - Scaffolding: the project
gatle-ignite initwrites trains, both from the source tree and from the installed wheel outside the repo, which proves the template ships in the wheel. - The examples run: every example except
mnist(it downloads) andteacher_student(it trains a teacher first) trains for one short epoch. That shows it runs, not that it learns: the scripts in an example'sscripts/check that, and are too slow to run on every commit. cfg.compileon CPU: the classification example trains through the CLI withcompile=True, a compiled model's checkpoint round-trips with an eager one, both refusals fire, and a 2-rank gloo run compiles and trains. No speed claim is made.
Verified once, on hardware¶
Weaker than the section above: each is one run, on one machine, on one day, and nothing re-runs it. It shows the code worked then, not that it still does.
- Multi-GPU: 12/12 checks on 2x Tesla T4, torch 2.10.0+cu128, via
scripts/verify_multigpu.py(nccl, exact eval sharding,sync_bn, bf16/fp16 autocast).
Not verified¶
- Multi-node (
nnodes > 1): no test has seen two nodes rendezvous. The topology maths is tested, and a run that would split into separate worlds is refused. Treat the first real run as the test. - Multi-config launch (
--config=a.py,b.py): the fan-out is tested, but not that each port it hands out is free.pick_portnever binds to check, so a clash with another process is possible. cfg.compileunder DDP on GPUs. A 2-rank gloo test compiles and trains on CPU, but never reaches NCCL or DDPOptimizer's bucketing.cfg.compileon a GPU:scripts/verify_compile.pyhas not been run on one.- External links in these docs are not checked.
Verification scripts¶
What the scripts in Verifying on real hardware have and have not run.
verify_multigpu.py: only check 1 (nccland exact sharding) has been seen to fail when what it checks is broken. Checks 2 (sync_bn) and 3 (autocast) need a GPU and have only run in the 12/12 pass above. If one fails, suspect the script before the framework.verify_compile.py: checks 6 and 7 skip, rather than pass, when the hardware is missing.- A wheel beside a checkout verifies the wheel. If they are different commits, the scripts test
the installed code. Say which you ran;
pip install -e .from the repo root avoids the question.
Known limitation¶
Resuming with a changed max_epochs restores a stale LR schedule. A scheduler's state_dict
carries its geometry (a cosine's T_max) as well as its position, so loading it overrides what the
config just computed. Train for 3 epochs, resume with max_epochs = 6, and the LR keeps following
the 3-epoch curve, with nothing to say so. The framework records max_epochs in each checkpoint and
warns when a resume disagrees with it.
To extend a run, don't resume. Start fresh from the weights: