Skip to content

fit.petab_v2.likelihood

The log-likelihood of an optimization problem.

A fit of sbmlsim is a weighted least squares fit, PEtab defines the objective of a problem as the likelihood of its measurements under a noise model. This module calculates that log-likelihood for the evaluation of a problem, the optimizer does not use it.

The noise model of a fit mapping is its sbmlsim.fit.objects.NoiseModel, i.e. the noise formula and the distribution of the observable it was read from. The simulation y is the median of the distribution of the measurement m and the noise formula gives its scale s, which are the definitions of PEtab v2:

normal        log p = -0.5 log(2 pi s^2)       - 0.5 ((m - y) / s)^2
log-normal    log p = -0.5 log(2 pi s^2 m^2)   - 0.5 ((log m - log y) / s)^2
laplace       log p = -log(2 s)                - |m - y| / s
log-laplace   log p = -log(2 s m)              - |log m - log y| / s

The functions of arrays (log_density, noise_values) know nothing of a problem, log_likelihood and gradient simulate one.

log_density

log_density(
    measurement, simulation, sigma, distribution=NORMAL
)

Get the log density of every measurement under its noise model.

Parameters:

Name Type Description Default
measurement ArrayLike

the measured values.

required
simulation ArrayLike

the simulated values at the measurements, the median of the distribution.

required
sigma ArrayLike

scale of the noise, one value or one value per measurement.

required
distribution NoiseDistribution

distribution of the noise.

NORMAL

Returns:

Type Description
ndarray

The log density of every measurement, see the module for the

ndarray

definitions.

Raises:

Type Description
ValueError

if the arrays do not have one shape, if a scale is not a positive finite number, if a measurement or a simulation is not finite, or if a measurement or a simulation of a logarithmic distribution is not positive.

noise_values

noise_values(noise, size, values=None, simulation=None)

Get the scale of the noise of every measurement of a fit mapping.

The symbols of the noise formula are resolved in this order: a placeholder is the value of the measurement, the symbol of the observable is the simulation, a parameter is the value of values and, without one, the nominal value of the noise model. A parameter which a problem estimates is therefore evaluated at the value it is given, it is not estimated.

Parameters:

Name Type Description Default
noise NoiseModel

noise model of the fit mapping.

required
size int

number of measurements.

required
values Mapping[str, float] | None

values of parameters by their id, e.g. the values of a parameter set.

None
simulation ArrayLike | None

the simulated values at the measurements, for a noise formula which holds the observable.

None

Returns:

Type Description
ndarray

The scale of the noise, one value per measurement.

Raises:

Type Description
ValueError

if the noise model has placeholders and not one value of them per measurement, if the formula holds the observable and no simulation is given, if a symbol has no value, or if a scale is not a positive finite number.

default_noise_model

default_noise_model(errors)

Get the noise model of a fit mapping which does not define one.

The noise is normal and its standard deviation is the error of the reference data as the problem resolves it (the SD, the SE without one, the largest error of the curve for a point without one), through the placeholder NOISE_PLACEHOLDER, and DEFAULT_SIGMA for data without errors. This is what the export writes for such a mapping, so the log-likelihood of a problem and of its PEtab problem agree.

Parameters:

Name Type Description Default
errors ArrayLike | None

the errors of the reference data of the mapping, i.e. OptimizationProblem.y_errors[k], None if it has none.

required

Returns:

Type Description
NoiseModel

The noise model.

noise_model_of

noise_model_of(problem, k)

Get the noise model of a fit mapping of an initialized problem.

Parameters:

Name Type Description Default
problem OptimizationProblem

initialized optimization problem.

required
k int

index of the fit mapping.

required

Returns:

Type Description
NoiseModel

The noise model of the mapping, default_noise_model if it has none.

nominal_parameters

nominal_parameters(problem)

Get the nominal values of the parameters of a problem.

The nominal value of a parameter is its start value, which is the nominalValue of the parameter table for a problem which was read, and the value of the model for a parameter without a start value.

Parameters:

Name Type Description Default
problem OptimizationProblem

initialized optimization problem.

required

Returns:

Type Description
ParameterSet

The parameter set nominal.

log_likelihood

log_likelihood(problem, parameters=None)

Get the log-likelihood of the training data of a problem.

The problem is simulated at the parameters and the log density of every measurement of the training data is summed, see log_density. The noise model of a fit mapping is noise_model_of. The settings of the fit, i.e. the residual, the weights and the loss function, do not enter.

Parameters:

Name Type Description Default
problem OptimizationProblem

initialized optimization problem.

required
parameters ParameterSet | None

parameters to evaluate the log-likelihood at, with the values of the parameters of the fit and, optionally, of parameters of the noise formulas. nominal_parameters by default.

None

Returns:

Type Description
float

The log-likelihood.

Raises:

Type Description
ValueError

if the problem is not initialized, if its residuals are relative to the baseline, if a simulation failed or if the noise model of a mapping cannot be evaluated.

KeyError

if the parameters lack a parameter of the fit.

stencil

stencil(
    value, h, lower_bound, upper_bound, order=2, name="?"
)

Get the points and the weights of the difference of a derivative.

The derivative is the sum of the weights times the values of the function at the points. The points stay inside the bounds and keep the step. The first difference which has all its points inside the bounds is taken:

order difference points error
4 central x -2h ... x +2h, four h^4
4 forward x ... x +4h, five h^4
4 backward x -4h ... x, five h^4
2 or 4 central x -h, x +h h^2
2 or 4 forward x ... x +2h, three h^2
2 or 4 backward x -2h ... x, three h^2
any secant of the bounds the bounds the distance of the bounds

A step which is shrunk to the distance to a bound, which is what a central difference next to a bound needs, divides the error of the function by a vanishing step, so the difference is one sided with the full step instead. The points are tested, not the distances to the bounds: a point which is one ulp outside a bound is not simulated.

The fall back to a lower order and the secant are logged as a warning, the derivative of a network next to a bound is then less exact than the order says.

Parameters:

Name Type Description Default
value float

the value of the parameter.

required
h float

the step.

required
lower_bound float

lower bound of the parameter.

required
upper_bound float

upper bound of the parameter.

required
order int

order of the central difference, 2 or 4.

2
name str

id of the parameter in the warnings.

'?'

Returns:

Type Description
list[tuple[float, float]]

The points with their weights.

Raises:

Type Description
ValueError

if the order is not 2 or 4, if the step is not positive, or if the value is outside the bounds or the bounds are equal.

gradient

gradient(
    problem, parameters=None, step=DEFAULT_STEP, order=2
)

Get the gradient of the log-likelihood by finite differences.

The differences are taken on the linear scale, i.e. in the units of the model, with the step step * max(|x|, 1) for a parameter of the value x, and are central differences, see stencil. The model is not simulated outside the bounds of a parameter: next to a bound the difference is one sided with the full step. A parameter without bounds has the central difference. The parameters whose difference falls back to fewer points or to the secant of the bounds are logged in one warning per fall back and call.

A difference divides the error of a simulation by the step, so the problem is initialized with FitSettings of tight tolerances and variable_step_size=False: with a variable step size the data is interpolated on the steps of the integrator, which differ between two simulations, and the simulations of one problem differ by 1e-6 however tight the tolerances are. The gradient logs a warning in this case.

The error of the central difference of three points grows with the square of the step and the third derivative. The log-likelihood of a model which oscillates is strongly curved in the parameters of a network: the difference of five points, order=4, has an error which grows with the fourth power of the step, for four simulations per parameter instead of two.

The gradient is defined inside the bounds only: a parameter outside its bounds raises, while log_likelihood evaluates the same parameters.

Parameters:

Name Type Description Default
problem OptimizationProblem

initialized optimization problem.

required
parameters ParameterSet | None

parameters to evaluate the gradient at, nominal_parameters by default.

None
step float

relative step of the differences.

DEFAULT_STEP
order int

order of the central difference, 2 or 4.

2

Returns:

Type Description
Series

The derivative of the log-likelihood by every parameter of the fit,

Series

indexed by the ids of the parameters.

Raises:

Type Description
ValueError

if the step is not positive, if the order is not 2 or 4, if a parameter is outside its bounds or its bounds are equal, or if the log-likelihood cannot be calculated, see log_likelihood.