On the dependency of neural differential equations on the training numerical method
Effects on optimisation and predictive performance
Publication date
2026-07-01
Document type
Konferenzbeitrag
Author
Organisational unit
Scopus ID
Conference
26th International Conference on Computational Science and Its Applications (ICCSA 2026) ; Braga, Portugal ; June 30 – July 3, 2026
Publisher
Springer Nature Switzerland
Book title
Computational Science and Its Applications – ICCSA 2026 Workshops
First page
255
Last page
271
Peer-reviewed
✅
Part of the university bibliography
✅
Language
English
Keyword
Continuous-normalizing flows
Differential Equations
Neural Networks
Neural ODE
Numerical Analysis
Numerical Methods
Optimisation
dtec.bw
Abstract
Neural Ordinary Differential Equations (Neural ODEs) model data as the solution of a parametrised ordinary differential equation, evaluated numerically at observation times. In practice, both training and prediction require the choice of a numerical method and its hyperparameters. While classical numerical analysis guarantees that all consistent methods converge to the same continuous solution as the step size tends to zero, Neural ODEs are trained and evaluated at finite resolution, where this choice becomes an integral part of the model itself. In this work, we theoretically show that the loss function optimised during training is a surrogate loss determined by the chosen numerical method and its configuration. Under standard regularity assumptions on the dynamics and the loss, we derive explicit bounds on how the order and step size of the integrator perturb the loss surface, its gradient, and the location of the resulting minimisers. These bounds quantify the discrepancy between training and testing losses when different solvers are used, and reveal that method-induced errors behave as a systematic model misspecification term. We validate these theory by providing numerical experiments on a popular spiral ODE benchmark, where we observe large train-test discrepancies and qualitatively different optimisation paths across solvers, even when training errors are nearly identical. Furthermore, the theoretical analysis developed in this work applies to any neural network architecture whose output is defined via a numerical solution of a differential equation, including Neural ODEs, continuous normalizing flows, and related models. This work calls for the attention to, and explicit reporting of, the numerical solvers and their configurations in the Neural ODE literature and in applications that rely on these models.
Version
Published version
Access right on openHSU
Metadata only access
