Skip to content

SBML

The Systems Biology Markup Language (SBML) is the format SBML4Humans reads. This page is the background a report assumes: what SBML is, how a model is built from its elements, what the Level 3 packages add, and how a model carries its metadata. The reference explains every element type and every attribute a report shows, one page per type; the sources of this page are listed under References.

What SBML is

SBML is a machine readable format for computational models of biological processes. It is written in XML, it belongs to no single program, and it is what lets a model which was built in one tool be simulated, analysed and published with another (Keating et al. 2020, doi:10.15252/msb.20199110). The alternative is a model which exists only in the file format of the program it was written in, or as equations printed in a paper, and neither can be executed by anyone else.

A model in SBML does not have to be written as a system of equations, and usually is not. It is written as biology: entities which are located in containers and are acted upon by processes (Keating et al. 2020). The entities are the species, the containers are the compartments, the processes are the reactions, and the equations a simulator integrates follow from that description. This is why the same file can be simulated, checked for consistency, drawn as a network and read by a person.

The format is defined by a specification document, not by an implementation. Everything a report of SBML4Humans shows is taken from that specification, which is why every reference page names the section it comes from.

Levels and versions

A level is a major edition of the language and represents a substantial change of its composition and structure; a version is a minor revision within a level which corrects, adjusts and refines its features (Hucka et al. 2019, doi:10.1515/jib-2019-0021). The levels stay distinct: all constructs of Level 1 can be mapped to Level 2 and all constructs of Level 2 to Level 3, but a valid Level 1 document is not a valid Level 2 document. A release is a corrected edition of one specification document and changes no syntax and no semantics.

SBML Level 3 Version 2 Core, Release 2 is the current core specification, and Level 3 is the only level which is modular, that is the only level with packages. Every file states the level and the version it is written in as attributes of its document, because they decide which constructs may appear in the file and how they are read. A report shows them in the context bar as L3V2 and in the attributes of the document.

Models of Level 1 and Level 2 are still common, for example in the curated models of BioModels. SBML4Humans reads a file in the level it is written in and does not convert it, so a report of a Level 2 model shows the elements that model actually has.

The structure of a model

A document contains at most one model, and the model contains the lists which everything else lives in. The model is also where the units of the whole model are declared, in particular the unit of time, which exists in no other element.

A compartment is a bounded space with a size, for example a cell, the cytosol or the plasma of an organism. A species is a pool of a chemical entity located in exactly one compartment, for example a metabolite, a protein or an ion; its quantity is an amount or a concentration and is what most simulations compute. A parameter is a named value the mathematics of the model can use, constant or changing over time.

A reaction is a process which changes the quantities of species. Its reactants and products are species references, which name a species and the stoichiometry it enters the reaction with, and its modifiers are modifier species references, which name a species that influences the reaction without being consumed, such as a catalyst or an inhibitor. How fast the reaction proceeds is given by its kinetic law, a formula which may use local parameters visible only inside that law. A reaction without a kinetic law is a structural statement, which is what a constraint based model is made of.

Not everything in a model is a reaction. An initial assignment computes the value of an element at the start of the simulation. An assignment rule states a formula which holds for its variable at every moment, a rate rule gives the rate of change of its variable, and an algebraic rule states an equation which has to hold at every moment without naming the variable it determines. A constraint states a condition a valid simulation has to satisfy; it changes nothing, but once it is violated the results from that moment on are no longer valid and the software has to say so.

An event is an instantaneous change of the model when a condition becomes true, for example a dose which is given at a fixed time. The condition is its trigger, what it changes is given by its event assignments, one per element it sets, and it may carry a delay, which is the time between the trigger and the change, and a priority, which orders it against the other events of the same moment. The trigger, the delay and the priority are elements of the file with a formula of their own, and the report shows them as such.

Two element types exist for the mathematics itself. A function definition is a named function which every formula of the model can call, so that a formula used in many reactions is written once. A unit definition is a named unit built from the base units of SBML with an exponent, a scale and a multiplier, for example millimole per litre, and it is what the units attributes of the other elements reference.

Every one of these elements carries the common attributes of SBase: the identifier other elements reference it by, a readable name, the meta id its annotations point at, a term of the Systems Biology Ontology, the notes for human readers and the annotations for machines.

Packages

A package of Level 3 adds elements and attributes for a domain which not every model needs. A file declares the packages it uses and whether understanding them is required to interpret the model, so that a tool knows what it is looking at instead of failing on unknown elements. This is what makes Level 3 an extensible format rather than one language which grows with every new need (Keating et al. 2020).

These are the packages with a published specification; further ones, among them arrays, dyn and spatial, exist as drafts.

package what it adds
comp Hierarchical Model Composition: a model is built out of other models, which are instantiated as submodels, trimmed by deletions and connected through ports and replacements
fbc Flux Balance Constraints: flux bounds, objective functions, gene associations and user defined constraints, that is what a constraint based model needs beyond the core
qual Qualitative Models: species which carry a level instead of an amount and transitions which decide that level, that is how a logical model or a Petri net is written
distrib Distributions: the uncertainty of a value and the distribution it was drawn from
groups Groups: a set of elements which belong together, for example the reactions of a pathway
layout Layout: the positions and the sizes of the elements in a diagram of the model
render Render: how the elements of a layout are drawn, with colours, strokes and gradients
multi Multistate, Multicomponent and Multicompartment Species: species with an internal state and components, and the rules which generate their reactions

Of these packages SBML4Humans reads comp, fbc, qual and distrib, which are the four the table links to their reference page. The submodels and the ports of comp, the gene products, the objectives, the flux bounds of fbc Version 1 and the user defined constraints of fbc Version 3, and the qualitative species and the transitions of qual become sections of the report like the types of the core. The external model definitions of comp belong to the document and are listed with it at the beginning of the type bar, and an uncertainty of distrib belongs to the element whose value it describes and is shown in the inspector of that element, with the uncert parameters and the uncert spans which are the measures it collects. The attributes these four packages add to the elements of the core are shown with the other attributes of the element.

A model which uses comp is built out of other models. A submodel instantiates a model definition of the same document or an external model definition of another file, a port is an element a model offers other models to connect to, a deletion removes an element of a submodel before it is instantiated, and a replaced element or a replaced by says which element of a submodel an element of the containing model takes the place of. Every one of them is an element of the report with its own attributes, and the report follows such a reference into the submodel, over a chain of references as deep as the file writes it, to the element it names. Where the model of the submodel is in another file, the report follows the reference into that file when it has it, which is the case when both files are entries of one COMBINE archive; Reading a report says what is read and what is not. Without the other file the reference keeps the name the file writes and has no link, which is how a reader tells a reference the report could follow from one it could not: a replacement still links the submodel it reaches into, and a deletion is named by its submodel, which lists it among its deletions.

The fbc package exists in three versions which differ in their data model, and the report reads what the version of the file defines. The flux bounds of Version 1 are objects of the model and become a section of the report; from Version 2 on the bounds of a reaction are two attributes which name a parameter, the model states whether it is strict, and the genes a reaction needs are a tree of and and or over gene product references, which the report shows as the expression that tree stands for, and across which a gene product names the reactions which need it. Version 3 adds the user defined constraints, which bound a weighted sum of fluxes and parameters rather than a single flux, a flux objective which may be quadratic, in one flux or in the product of two, and the key value pairs any element may carry.

The distrib package records how well a value is known. An uncertainty belongs to the element whose value it describes and collects the measures of that value: an uncert parameter is one measure, a mean, a standard deviation, a sample size or a distribution with the parameters it is defined by, and an uncert span is a measure which is an interval, such as a range or a confidence interval, whose ends are numbers or elements of the model. The report shows the measures of an element in the inspector of that element, an interval as the interval it is, the uncertainty leads back to the element it describes, and a measure carries the notes and the annotations in which a file records which experiment or which publication a number comes from.

A qualitative model is the one kind of model in this list which is not built from reactions at all. Its entities carry a level, a whole number which stands for a range of activity such as "off" and "on" rather than a concentration, and a transition says which level an entity takes next given the levels of the entities which act on it. That is how a regulatory network is written down when its kinetics are unknown, which is the usual case for a network of genes, and the report shows it the way such a model is read: the qualitative species with their levels, the inputs of a transition with the sign which says whether they activate or inhibit, and the function terms with the default term behind them as the transition table they are.

The packages a file declares are shown in the bar at the top of the report, whether the report reads them or not, and the type bar carries an element type of a package only when the file declares it and the model states elements of that type. What another package adds is not lost: it stays in the XML of the element which carries it, which the inspector shows.

Annotations

A model which is only correct is not yet reusable: a species called x1 says nothing about which molecule it is. SBML therefore gives every element two places for metadata, and the report shows both.

The notes of an element are XHTML written for human readers, for example the derivation of a rate law or the source of a parameter value. The annotation of an element is RDF written for machines and follows the MIRIAM guidelines for the annotation of biochemical models (Le Novère et al. 2005, doi:10.1038/nbt1156).

A MIRIAM style annotation is a set of controlled vocabulary terms, each of which is a qualifier and the resources it relates the element to. The qualifier states the relation and comes from one of two BioModels.net namespaces: a biological qualifier such as bqbiol:is, bqbiol:hasPart or bqbiol:isVersionOf relates the biological entity the element stands for to the resource, a model qualifier such as bqmodel:is or bqmodel:isDerivedFrom relates the model itself to it. The distinction matters: bqbiol:is on a species says the species is that molecule, bqmodel:is on a model says the model is that entry in a model database.

A resource is a URI which identifies an entry of a database, usually of the form https://identifiers.org/<collection>/<identifier>, for example https://identifiers.org/uniprot/P12999 or https://identifiers.org/chebi/CHEBI:17234. Older files write the same reference as a MIRIAM URN, for example urn:miriam:kegg.compound:C00046, which the repressilator example of the application uses; the report links such a URN to the same entry. The registry behind identifiers.org resolves the identifier to the databases which hold the entry. The report asks its backend to do that by itself for at most a hundred resources of an element, spent over the terms it shows, of which there are fifty before a click, so that the annotations of an element show what an entry is called instead of the identifier alone. Everything beyond that budget is resolved on a click.

Independently of these annotations, every element may carry a term of the Systems Biology Ontology in its sbo attribute. The ontology names what an element is in the vocabulary of systems biology, for example SBO:0000247 for a simple chemical or SBO:0000185 for a transport reaction, which is a statement about the role of the element in the model rather than about the molecule it represents. The report links the term to its entry and also lists it among the annotations of the element.

Besides the qualifiers, the RDF annotation carries the history of the SBML encoding, in elements of its own which stand in front of them: who created it, with which organization and mail address, when it was created and when it was modified. It is the history of the encoding, not of the model the encoding stands for. The report shows it below the notes of an element.

COMBINE archives

A model is rarely the whole story of a study: there are the simulation descriptions, the data, the figures and the documentation next to it. A COMBINE archive is the standard container for these files. It is a zip file with a manifest which lists every entry with its location and its format, and marks the entries a reader should start with as master entries. Its usual file extension is .omex.

SBML4Humans creates one report per SBML entry of an archive and offers the entries in the context bar, starting with the master entry if it has a report and with the first entry otherwise, because an archive does not have to mark one. An SBML file which is submitted on its own is wrapped in an archive with a single master entry, so that every report comes with a manifest and an archive and a plain file are read the same way.