Study format¶
Study format 2 is being introduced
This page describes study format 2. The studies in pkdb_data are still in study format 1 until the migration; pkdb format, pkdb validate, pkdb prepare and pkdb upload all support study format 2 folders.
A study folder in study format 2 contains study.json, reference.json, review.json, the tab-separated tables below, the publication PDF and an image <study>_<source>.png for every paper table or figure the data come from. pkdb format writes all files in canonical form, pkdb validate checks them, and pkdb schema export writes the same definitions as JSON Schema.
Identity and upload¶
A study is identified by its folder location <substance>/<name>, for example caffeine/Harder1988: the folder name is the study name and its parent folder the substance. The folder names publication and validate are reserved, because PK-DB uses them in study URLs. pkdb upload sends the exact text of study.json and reference.json and every other file of the folder to /api/v2/studies/<substance>/<name>, where the server validates the folder like pkdb validate and stores the study. The study page is /data/<substance>/<name>.
A released study has a PKDB identifier in the release block of study.json. Uploading a folder with a release.pkdb_id takes over the study that PK-DB stores under that identifier and renames it to the folder location, so a moved or renamed folder keeps its study. A folder without one, or whose PKDB identifier no stored study has, takes over the study format 1 study of the same publication and source, if you may edit it. PKDB identifiers and former study format 1 identifiers in URLs, such as /data/PKDB00198 or /data/Vilsboll2008, redirect to the study.
Statistics and schedules¶
There is no value column. mean is the arithmetic mean, the value of a single subject (count 1) or, with the calculation unspecified summary, the central value of a summary whose statistic the publication does not state. cv and gcv are entered in percent and are fractions in the prepared study and the API. A digitized error_bar with its error_type completes sd, se or gsd; missing sd, se and cv are derived from each other, and gsd and gcv as well, where the reported values allow it, and reported statistics that contradict each other, including an error bar, are a warning that a review item acknowledges at the column of the statistic that disagrees with the most others (see the data model). An intervention time is a number, or a ;-separated list for an irregular schedule; a regular schedule uses interval and doses instead.
Tables¶
| File | Content | Required |
|---|---|---|
subjects.tsv |
Groups and individuals. A subject with count 1 is an individual; every subject except all has a parent group. |
yes |
interventions.tsv |
Interventions such as doses, fasting or smoking, referenced by name from the other tables. | no |
characteristica.tsv |
Baseline values of subjects, such as age, weight, sex or creatinine clearance. | no |
outputs_<source>.tsv |
Single values after interventions, such as pharmacokinetic parameters, from one paper table or figure. | no |
timecourses_<source>.tsv |
Timecourse points from one paper figure or table. Rows with the same label form one series. | no |
scatters_<source>.tsv |
Points of scatter plots, one row per point with an x and a y value. | no |
TSV encoding¶
- UTF-8 without byte order mark, LF line endings and a final newline.
- Tab separated without quoting. Cells contain no tabs or line breaks.
- One header row with every template column in template order.
- An empty cell means missing.
NR(not reported) is allowed only intimeandtime_unitand their scatter variantsx_time,x_time_unit,y_timeandy_time_unit. - Numbers use a decimal point and are written in their shortest form.
pkdb formatsorts the rows and writes thestudycolumn and, in files named after a source, thesourcecolumn.
subjects.tsv¶
Groups and individuals. A subject with count 1 is an individual; every subject except all has a parent group.
| Column | Type | Required | Description |
|---|---|---|---|
study |
text | no, written by pkdb format |
Study name, written by pkdb format from the folder name. Do not edit. |
name |
name | yes | Unique subject name. The root group is all. |
parent |
name | no | Name of the parent group. Empty only for all. |
count |
whole number | no | Number of subjects. A count of 1 makes the row an individual; any other count, or an empty cell, makes it a group. |
source |
source | no | Paper table or figure the row comes from, such as Tab1 or Fig2A, or Text for the article text. Except for Text, the image <study>_<source>.png must exist. |
comment |
text | no | Free-text comment. |
interventions.tsv¶
Interventions such as doses, fasting or smoking, referenced by name from the other tables.
| Column | Type | Required | Description |
|---|---|---|---|
study |
text | no, written by pkdb format |
Study name, written by pkdb format from the folder name. Do not edit. |
source |
source | no | Paper table or figure the row comes from, such as Tab1 or Fig2A, or Text for the article text. Except for Text, the image <study>_<source>.png must exist. |
name |
name | yes | Unique intervention name, referenced from the interventions columns of other tables. |
subjects |
name | no | Subject that the dose statistics describe, for body-weight-adjusted doses with sd, min or max. Usually empty. |
measurement |
vocabulary: measurements | yes | Kind of intervention from the vocabulary, such as dosing, qualitative dosing or fasting. |
calculation |
vocabulary: calculation types | no | How the central value was obtained, from the vocabulary. unspecified summary marks a central value whose statistic the publication does not state; enter it in mean. |
substance |
vocabulary: substances | no | Substance from the vocabulary. |
tissue |
vocabulary: tissues | no | Tissue or matrix from the vocabulary. |
method |
vocabulary: methods | no | Analytical method from the vocabulary. |
choice |
text | no | Categorical value of a categorical intervention, such as Y for fasting. |
route |
vocabulary: routes | no | Administration route from the vocabulary. |
form |
vocabulary: forms | no | Administration form from the vocabulary. |
application |
vocabulary: applications | no | Application from the vocabulary, such as single dose. |
time |
number, ;-separated numbers or NR |
no | Time of the first administration in time_unit. An irregular schedule is a ;-separated list such as 0;12;40. |
time_end |
number | no | End of a continuous administration in time_unit. |
interval |
number | no | Dosing interval in time_unit. |
doses |
whole number | no | Number of administrations. |
time_unit |
unit or NR |
no | Unit of time, time_end and interval, or NR when the publication does not report it. |
count |
whole number | no | Number of subjects the dose statistics describe. Usually empty. |
mean |
number | no | Dose. A fixed dose is entered as its value, such as 100 with unit mg. |
sd |
number | no | Standard deviation. |
se |
number | no | Standard error of the mean. |
cv |
number | no | Coefficient of variation in percent. |
gmean |
number | no | Geometric mean. |
gsd |
number | no | Geometric standard deviation as a dimensionless factor of at least 1. |
gcv |
number | no | Geometric coefficient of variation in percent. |
median |
number | no | Median. |
min |
number | no | Minimum. |
max |
number | no | Maximum. |
unit |
unit | no | Unit of mean, sd, se, gmean, median, min, max and error_bar. |
error_bar |
number | no | Digitized end of an error bar on the value axis, in unit. The statistic named in error_type is derived from it. |
error_type |
one of sd, se, gsd |
no | Statistic the error bar shows: sd, se or gsd. |
comment |
text | no | Free-text comment. |
characteristica.tsv¶
Baseline values of subjects, such as age, weight, sex or creatinine clearance.
| Column | Type | Required | Description |
|---|---|---|---|
study |
text | no, written by pkdb format |
Study name, written by pkdb format from the folder name. Do not edit. |
source |
source | yes | Paper table or figure the row comes from, such as Tab1 or Fig2A, or Text for the article text. Except for Text, the image <study>_<source>.png must exist. |
subjects |
name | yes | Name of the subject in subjects.tsv that the row describes: a group, or an individual with count 1. |
measurement |
vocabulary: measurements | yes | Baseline quantity from the vocabulary, such as age, weight or sex. |
calculation |
vocabulary: calculation types | no | How the central value was obtained, from the vocabulary. unspecified summary marks a central value whose statistic the publication does not state; enter it in mean. |
substance |
vocabulary: substances | no | Substance from the vocabulary. |
tissue |
vocabulary: tissues | no | Tissue or matrix from the vocabulary. |
method |
vocabulary: methods | no | Analytical method from the vocabulary. |
choice |
text | no | Categorical value of a categorical measurement, such as M for sex. |
time |
number or NR |
no | Time point in time_unit, or NR when the publication does not report it. |
time_unit |
unit or NR |
no | Unit of time, or NR when the publication does not report it. |
count |
whole number | no | Number of subjects the row describes. Leave empty when it equals the count of the referenced subject. In a choice row, the number of subjects with that choice. |
mean |
number | no | Arithmetic mean. For a single subject (count 1) and for unspecified summary, the reported value. |
sd |
number | no | Standard deviation. |
se |
number | no | Standard error of the mean. |
cv |
number | no | Coefficient of variation in percent. |
gmean |
number | no | Geometric mean. |
gsd |
number | no | Geometric standard deviation as a dimensionless factor of at least 1. |
gcv |
number | no | Geometric coefficient of variation in percent. |
median |
number | no | Median. |
min |
number | no | Minimum. |
max |
number | no | Maximum. |
unit |
unit | no | Unit of mean, sd, se, gmean, median, min, max and error_bar. |
error_bar |
number | no | Digitized end of an error bar on the value axis, in unit. The statistic named in error_type is derived from it. |
error_type |
one of sd, se, gsd |
no | Statistic the error bar shows: sd, se or gsd. |
comment |
text | no | Free-text comment. |
outputs_<source>.tsv¶
Single values after interventions, such as pharmacokinetic parameters, from one paper table or figure.
| Column | Type | Required | Description |
|---|---|---|---|
study |
text | no, written by pkdb format |
Study name, written by pkdb format from the folder name. Do not edit. |
source |
source | no, written by pkdb format |
Paper table or figure of the row, written by pkdb format from the file name. Do not edit. |
subjects |
name | yes | Name of the subject in subjects.tsv that the row describes: a group, or an individual with count 1. |
interventions |
comma-separated names | no | Comma-separated names of the interventions in interventions.tsv that the subjects received before the measurement. Empty for none. |
measurement |
vocabulary: measurements | yes | Measured quantity from the vocabulary, such as cmax, auc_inf or thalf. |
calculation |
vocabulary: calculation types | no | How the central value was obtained, from the vocabulary. unspecified summary marks a central value whose statistic the publication does not state; enter it in mean. |
substance |
vocabulary: substances | no | Substance from the vocabulary. |
tissue |
vocabulary: tissues | no | Tissue or matrix from the vocabulary. |
method |
vocabulary: methods | no | Analytical method from the vocabulary. |
choice |
text | no | Categorical value of a categorical measurement, such as M for sex. |
time |
number or NR |
no | Time point in time_unit, or NR when the publication does not report it. |
time_unit |
unit or NR |
no | Unit of time, or NR when the publication does not report it. |
count |
whole number | no | Number of subjects the row describes. Leave empty when it equals the count of the referenced subject. In a choice row, the number of subjects with that choice. |
mean |
number | no | Arithmetic mean. For a single subject (count 1) and for unspecified summary, the reported value. |
sd |
number | no | Standard deviation. |
se |
number | no | Standard error of the mean. |
cv |
number | no | Coefficient of variation in percent. |
gmean |
number | no | Geometric mean. |
gsd |
number | no | Geometric standard deviation as a dimensionless factor of at least 1. |
gcv |
number | no | Geometric coefficient of variation in percent. |
median |
number | no | Median. |
min |
number | no | Minimum. |
max |
number | no | Maximum. |
unit |
unit | no | Unit of mean, sd, se, gmean, median, min, max and error_bar. |
error_bar |
number | no | Digitized end of an error bar on the value axis, in unit. The statistic named in error_type is derived from it. |
error_type |
one of sd, se, gsd |
no | Statistic the error bar shows: sd, se or gsd. |
comment |
text | no | Free-text comment. |
timecourses_<source>.tsv¶
Timecourse points from one paper figure or table. Rows with the same label form one series.
| Column | Type | Required | Description |
|---|---|---|---|
study |
text | no, written by pkdb format |
Study name, written by pkdb format from the folder name. Do not edit. |
source |
source | no, written by pkdb format |
Paper table or figure of the row, written by pkdb format from the file name. Do not edit. |
label |
name | yes | Name of the timecourse. Rows with the same label form one series; labels are unique across the study. |
subjects |
name | yes | Name of the subject in subjects.tsv that the row describes: a group, or an individual with count 1. |
interventions |
comma-separated names | no | Comma-separated names of the interventions in interventions.tsv that the subjects received before the measurement. Empty for none. |
measurement |
vocabulary: measurements | yes | Measured quantity from the vocabulary, such as concentration. |
calculation |
vocabulary: calculation types | no | How the central value was obtained, from the vocabulary. unspecified summary marks a central value whose statistic the publication does not state; enter it in mean. |
substance |
vocabulary: substances | no | Substance from the vocabulary. |
tissue |
vocabulary: tissues | no | Tissue or matrix from the vocabulary. |
method |
vocabulary: methods | no | Analytical method from the vocabulary. |
choice |
text | no | Categorical value of a categorical measurement, such as M for sex. |
time |
number or NR |
yes | Time point in time_unit, or NR when the publication does not report it. |
time_unit |
unit or NR |
yes | Unit of time, or NR when the publication does not report it. |
count |
whole number | no | Number of subjects the row describes. Leave empty when it equals the count of the referenced subject. In a choice row, the number of subjects with that choice. |
mean |
number | no | Arithmetic mean. For a single subject (count 1) and for unspecified summary, the reported value. |
sd |
number | no | Standard deviation. |
se |
number | no | Standard error of the mean. |
cv |
number | no | Coefficient of variation in percent. |
gmean |
number | no | Geometric mean. |
gsd |
number | no | Geometric standard deviation as a dimensionless factor of at least 1. |
gcv |
number | no | Geometric coefficient of variation in percent. |
median |
number | no | Median. |
min |
number | no | Minimum. |
max |
number | no | Maximum. |
unit |
unit | no | Unit of mean, sd, se, gmean, median, min, max and error_bar. |
error_bar |
number | no | Digitized end of an error bar on the value axis, in unit. The statistic named in error_type is derived from it. |
error_type |
one of sd, se, gsd |
no | Statistic the error bar shows: sd, se or gsd. |
comment |
text | no | Free-text comment. |
scatters_<source>.tsv¶
Points of scatter plots, one row per point with an x and a y value.
| Column | Type | Required | Description |
|---|---|---|---|
study |
text | no, written by pkdb format |
Study name, written by pkdb format from the folder name. Do not edit. |
source |
source | no, written by pkdb format |
Paper table or figure of the row, written by pkdb format from the file name. Do not edit. |
name |
name | yes | Name of the scatter dataset, unique across the study. |
subjects |
name | yes | Subject of the point, usually an individual with count 1. The points of one scatter are all individuals or all groups. |
x_interventions |
comma-separated names | no | X axis: comma-separated names of the interventions in interventions.tsv that the subjects received before the measurement. Empty for none. |
x_measurement |
vocabulary: measurements | yes | X axis: measured quantity from the vocabulary, such as cmax, age or weight. |
x_substance |
vocabulary: substances | no | X axis: substance from the vocabulary. |
x_tissue |
vocabulary: tissues | no | X axis: tissue or matrix from the vocabulary. |
x_method |
vocabulary: methods | no | X axis: analytical method from the vocabulary. |
x_time |
number or NR |
no | X axis: time point in x_time_unit, or NR when the publication does not report it. |
x_time_unit |
unit or NR |
no | X axis: unit of x_time, or NR when the publication does not report it. |
x_mean |
number | yes | X axis: value of the point. |
x_unit |
unit | no | X axis: unit of the value. |
y_interventions |
comma-separated names | no | Y axis: comma-separated names of the interventions in interventions.tsv that the subjects received before the measurement. Empty for none. |
y_measurement |
vocabulary: measurements | yes | Y axis: measured quantity from the vocabulary, such as cmax, age or weight. |
y_substance |
vocabulary: substances | no | Y axis: substance from the vocabulary. |
y_tissue |
vocabulary: tissues | no | Y axis: tissue or matrix from the vocabulary. |
y_method |
vocabulary: methods | no | Y axis: analytical method from the vocabulary. |
y_time |
number or NR |
no | Y axis: time point in y_time_unit, or NR when the publication does not report it. |
y_time_unit |
unit or NR |
no | Y axis: unit of y_time, or NR when the publication does not report it. |
y_mean |
number | yes | Y axis: value of the point. |
y_unit |
unit | no | Y axis: unit of the value. |
comment |
text | no | Free-text comment. |