# pkpdbib > Python utilities for PK/PD literature and bibliography management The complete documentation from https://matthiaskoenig.github.io/pkpdbib, one section per page. --- # pkpdbib ![pkpdbib logo](images/favicon/pkpdbib-100x100-300dpi.png){ align=left width=100 } `pkpdbib` is a collection of python utilities for working with pharmacokinetics/pharmacodynamics (PK/PD) literature and bibliographies. It supports the curation of the literature of [PK-DB](https://pk-db.com) and of the PK/PD models built on it. Features include - **[PDF retrieval](pdfs.md)** - the PDFs of the DOIs of a [Zotero](https://www.zotero.org) export (Better BibTeX JSON), via the `scihub_pdfs` command - **[Zotero libraries](zotero.md)** - programmatic access to a Zotero library and a table of the tags of its items The source code is available from [github.com/matthiaskoenig/pkpdbib](https://github.com/matthiaskoenig/pkpdbib). If you have any questions or issues please [open an issue](https://github.com/matthiaskoenig/pkpdbib/issues). ## Quickstart ```bash pip install pkpdbib scidownl domain.update scihub_pdfs -j aliskiren.json ``` See [Installation](installation.md) and [PDF retrieval](pdfs.md) for the details. ## How to cite [![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.11076700.svg)](https://doi.org/10.5281/zenodo.11076700) If you use `pkpdbib` please cite the archived software on [Zenodo](https://doi.org/10.5281/zenodo.11076700). The DOI always resolves to the latest version, the current release is archived as: > König, M. (2026). *pkpdbib: python utilities for PK/PD literature and bibliography management* (Version 0.2.1) [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.23154802 ```bibtex @software{konig_pkpdbib, author = {König, Matthias}, title = {pkpdbib: python utilities for PK/PD literature and bibliography management}, year = {2026}, month = oct, version = {0.2.1}, publisher = {Zenodo}, doi = {10.5281/zenodo.23154802}, url = {https://doi.org/10.5281/zenodo.23154802}, } ``` The citation metadata is in [CITATION.cff](https://github.com/matthiaskoenig/pkpdbib/blob/develop/CITATION.cff). ## License - Source code: [MIT](https://opensource.org/license/MIT) - Documentation: [CC BY-SA 4.0](http://creativecommons.org/licenses/by-sa/4.0/) ## Funding Matthias König is supported by the German Research Foundation (DFG) within the Research Unit Programme FOR 5151 "QuaLiPerF (Quantifying Liver Perfusion-Function Relationship in Complex Resection - A Systems Medicine Approach)" by grant number 436883643 and by grant number 465194077 (Priority Programme SPP 2311, Subproject SimLivA). Matthias König was supported by the Federal Ministry of Education and Research (BMBF, Germany) within the research network Systems Medicine of the Liver (**LiSyM**, grant number 031L0054) and within ATLAS by grant number 031L0304B. --- # Installation `pkpdbib` requires python >= 3.13 and is available from [pypi](https://pypi.org/project/pkpdbib/). It is tested on Linux, macOS and Windows. ## With uv [uv](https://docs.astral.sh/uv/) is the recommended way to install the package. In a project it is added as a dependency: ```bash uv add pkpdbib ``` The `scihub_pdfs` command can also be used without installing the package into an environment: ```bash uvx --from pkpdbib scihub_pdfs -j aliskiren.json ``` ## With pip ```bash pip install pkpdbib ``` ## Dependencies `pkpdbib` depends on `rich` for the console output, `scidownl` for the [PDF retrieval](pdfs.md), and `pyzotero` and `polars` for the [Zotero libraries](zotero.md). ## Development version The current state of the `develop` branch is installed directly from GitHub: ```bash uv add "pkpdbib @ git+https://github.com/matthiaskoenig/pkpdbib.git@develop" ``` or, with pip, ```bash pip install git+https://github.com/matthiaskoenig/pkpdbib.git@develop ``` To work on the repository itself, with the test and documentation tooling, see [Development](development.md). --- # PDF retrieval `pkpdbib.scihub_tools` downloads the PDFs of the DOIs of a [Zotero](https://www.zotero.org) export via [scidownl](https://github.com/Tishacy/SciDownl). The export uses the Better BibTeX JSON format of the [Better BibTeX](https://retorque.re/zotero-better-bibtex/) plugin, the PDFs are named by the citation keys of the items. !!! warning Only download publications you are entitled to access. Check the copyright law of your country and the terms of your institution before using Sci-Hub. ## Export from Zotero 1. Open the Zotero library and select the items without a PDF attachment. 2. Right click -> `Export Items...` -> format `Better BibTeX JSON`. 3. Save the export as `.json`, e.g., `aliskiren.json`. ## Command line Update the list of the available Sci-Hub domains once, then download the PDFs: ```bash scidownl domain.update scihub_pdfs -j aliskiren.json ``` The PDFs are written to the directory `aliskiren/` next to the JSON file, e.g., `aliskiren/limoges2008.pdf`. PDFs which exist already are not downloaded again, so the command can be repeated after a partial run. The DOIs for which no PDF could be retrieved are listed at the end. | option | description | | --- | --- | | `-j`, `--json` | path of the Better BibTeX JSON file, required | | `-o`, `--out` | directory of the PDFs, by default named after the JSON file | | `--scihub-url` | Sci-Hub domain to use, by default the best domain of scidownl | ## Python ```python from pathlib import Path from pkpdbib.scihub_tools import dois_from_json, scihub_pdfs # DOIs by citation key, items without DOI are skipped dois = dois_from_json(Path("aliskiren.json")) # download the missing PDFs, returns the keys without PDF missing = scihub_pdfs(Path("aliskiren.json")) ``` See the [API reference](api/scihub_tools.md) for all functions. --- # Zotero libraries `pkpdbib.zotero_tools` provides programmatic access to [Zotero](https://www.zotero.org) libraries via [pyzotero](https://pyzotero.readthedocs.io) and creates tables of the tags of their items with [polars](https://pola.rs). ## Access Access needs an API key with read access to the library, created at [zotero.org/settings/keys/new](https://www.zotero.org/settings/keys/new). The id of a group library is found by opening the page of the group via [zotero.org/groups](https://www.zotero.org/groups) and hovering over the link to the group settings, it is the integer after `/groups/`. Keep the API key out of the source code, e.g., in an environment variable: ```python import os from pkpdbib.zotero_tools import create_zot_client, get_items zot = create_zot_client( api_key=os.environ["ZOTERO_API_KEY"], library_id=4979949, library_type="group", ) items = get_items(zot) # all top level items ``` ## Tag tables The curation of a library is recorded in the tags of its items, e.g., `pkdb` for items curated in [PK-DB](https://pk-db.com) or `species:human`. `create_tag_table` creates a table with one row per item and one boolean column per selected tag, in addition to the DOI and the PubMed id of the item: ```python from pkpdbib.zotero_tools import create_tag_table df = create_tag_table( items, tags_set={"pkdb", "has_simulation"}, tag_prefixes=["species:", "data:"], ) df.write_csv("tags.csv") ``` A complete example is [`examples/zotero_tags.py`](https://github.com/matthiaskoenig/pkpdbib/blob/develop/examples/zotero_tags.py). See the [API reference](api/zotero_tools.md) for all functions. --- # API reference The API reference is generated from the docstrings of the package. | module | description | | --- | --- | | [scihub_tools](scihub_tools.md) | the PDFs of the DOIs of a Zotero export, the `scihub_pdfs` command | | [zotero_tools](zotero_tools.md) | access to Zotero libraries and tables of their tags | | [console](console.md) | shared rich console | | [log](log.md) | logging of the package | --- # pkpdbib.scihub_tools Retrieve the PDFs of DOIs via Sci-Hub. The DOIs are read from a Zotero export in the Better BibTeX JSON format (`Export Items...` -> `Better BibTeX JSON`), the PDFs are named by the citation keys of the items. ## function `dois_from_json(json_path: pathlib.Path) -> dict[str, str]` Read the DOIs of a Better BibTeX JSON export of Zotero. Items without a DOI are skipped. Args: json_path: path of the Better BibTeX JSON file. Returns: The DOIs by the citation keys of the items. ## function `scihub_pdf_from_doi(doi: str, pdf_path: pathlib.Path, scihub_url: str | None = None) -> None` Download the PDF of a DOI. Args: doi: DOI of the publication, e.g., `10.1145/3375633`. pdf_path: path the PDF is written to. scihub_url: Sci-Hub domain to use; by default the best available domain of scidownl, see `scidownl domain.update`. ## function `scihub_pdfs(zotero_json_path: pathlib.Path, pdf_dir: pathlib.Path | None = None, scihub_url: str | None = None) -> list[str]` Download the missing PDFs of a Better BibTeX JSON export of Zotero. Args: zotero_json_path: path of the Better BibTeX JSON file. pdf_dir: directory of the PDFs; by default the directory named after the JSON file next to it, i.e., `/` for `.json`. scihub_url: Sci-Hub domain to use, see `scihub_pdf_from_doi`. Returns: The citation keys for which no PDF exists after the download. ## function `scihub_pdfs_command(argv: collections.abc.Sequence[str] | None = None) -> None` Command line interface `scihub_pdfs`, see `scihub_pdfs`. Args: argv: command line arguments, by default `sys.argv[1:]`. ## function `scihub_pdfs_from_dois(dois: dict[str, str], pdf_dir: pathlib.Path, scihub_url: str | None = None) -> list[str]` Download the PDFs of DOIs which are not in the directory yet. Args: dois: DOIs by the keys the PDFs are named after, `.pdf`. pdf_dir: directory of the PDFs, created if it does not exist. scihub_url: Sci-Hub domain to use, see `scihub_pdf_from_doi`. Returns: The keys for which no PDF exists after the download. --- # pkpdbib.zotero_tools Programmatic access to Zotero libraries. - documentation of pyzotero: https://pyzotero.readthedocs.io - API key: https://www.zotero.org/settings/keys/new - group id: open the page of the group via https://www.zotero.org/groups and hover over the link to the group settings, the id is the integer after `/groups/`. ## function `create_tag_table(items: collections.abc.Iterable[dict[str, typing.Any]], tags_set: collections.abc.Iterable[str], tag_prefixes: collections.abc.Iterable[str]) -> polars.dataframe.frame.DataFrame` Create a table of the items with the selected tags. A tag is selected if it is in `tags_set` or starts with one of the `tag_prefixes`. Every selected tag is a boolean column, which is `True` for the items with the tag. Args: items: items of a library, see `get_items`. tags_set: tags to select. tag_prefixes: prefixes of the tags to select, e.g., `"species:"`. Returns: The table with the columns `key`, `doi`, `pubmed` and one column per selected tag. ## function `create_zot_client(api_key: str, library_id: int, library_type: str = 'group') -> pyzotero._client.Zotero` Create a Zotero client bound to a library. The item methods of the client only operate on this library. Args: api_key: Zotero API key with read access to the library. library_id: id of the user or group library. library_type: `group` or `user`. Returns: The client. ## function `get_items(zot: pyzotero._client.Zotero, show: bool = False, limit: int | None = None) -> list[dict[str, typing.Any]]` Get the top level items of the library. Args: zot: client of the library, see `create_zot_client`. show: print the items to the console. limit: maximal number of items; all items if `None`. Returns: The items as returned by the Zotero API. ## function `pmid_from_extra(extra: str | None) -> str | None` Get the PubMed id from the `extra` field of a Zotero item. Args: extra: `extra` field with one `key: value` per line, e.g., `PMID: 27267043` and `PMCID: PMC4895977`. Returns: The PubMed id or `None`. --- # pkpdbib.console Shared rich console of the package. --- # pkpdbib.log Logging of the package, rendered with rich. ## function `get_logger(name: str, level: int = 20) -> logging.Logger` Get the logger for the given name, logging via the shared console. The rich handler is added once, so repeated calls for the same name do not duplicate the output. Args: name: name of the logger, typically `__name__` of the module. level: logging level of the logger. Returns: The configured logger. --- # Development Contributions are welcome. The repository is [matthiaskoenig/pkpdbib](https://github.com/matthiaskoenig/pkpdbib); development happens against the `develop` branch via pull requests. ## Branch model Two branches are permanent: - **`develop`** is the default branch and the branch everything is integrated into. The documentation on [matthiaskoenig.github.io/pkpdbib](https://matthiaskoenig.github.io/pkpdbib) is published from it. - **`main`** tracks the latest published release. It is fast-forwarded to the released commit by the `sync-main` job of the `CI-CD` workflow after the package went to pypi, so `main` and the newest version on pypi always agree. Nothing is developed on `main` and nothing is merged into it by hand. Work happens on short lived branches off `develop`, which GitHub deletes after the merge. Releases are tagged on `develop`, see [Release](#release). ## Pull requests Neither branch accepts a direct push, every change goes through a pull request against `develop`. This includes the maintainer, there is no bypass. A pull request can only be merged once the four required checks are green: | check | workflow | content | | ------- | ------------- | -------------------------------------------------------------------- | | `tests` | `ci-cd.yml` | the test matrix, linux with python 3.13 to 3.15, macos and windows with 3.15, and the lower bounds of the dependencies (`lowest`) | | `ruff` | `ruff.yml` | `ruff check` and `ruff format --check` | | `ty` | `ty.yml` | `tox r -e ty` | | `docs` | `docs.yml` | the zensical build including the api reference and the agent files | `tests` aggregates the test matrix into a single job, so the name of the required check stays the same when the matrix changes. Further rules of a pull request: - conversations have to be resolved before the merge - an approval is dismissed when new commits are pushed - the history stays linear, i.e., a pull request is merged with squash or rebase; merge commits are disabled - the maintainer is the code owner of the repository (`.github/CODEOWNERS`) and is requested for review on every pull request. A pull request of a contributor is therefore reviewed and merged by the maintainer, who has the only write access. The rulesets themselves do not require an approval: on a personal repository a ruleset cannot ask for an approval only from somebody else, and requiring one would block the pull requests of the maintainer, who cannot approve their own. Once a second person has write access, a ruleset requiring an approving review of a code owner can be added [Auto-merge](https://docs.github.com/pull-requests/collaborating-with-pull-requests/incorporating-changes-from-a-pull-request/automatically-merging-a-pull-request) is enabled for the repository, so a pull request can be queued and is merged as soon as the checks pass and the required approval is there. ### Repository policies { #repository-policies } The protection is implemented with [repository rulesets](https://docs.github.com/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/about-rulesets). They are part of the repository in `.github/rulesets/` instead of only living in the web interface, so a change to a policy is reviewed like any other change: | ruleset | applies to | rules | | ----------------------- | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------- | | `develop.json` | `develop` | pull request required, the four checks above, resolved conversations, linear history, no force push, no deletion. **No bypass, for anybody.** | | `main.json` | `main` | linear history, no force push, no deletion, no bypass. The fast-forward of the release workflow needs none, only a force push or a merge commit would be rejected | | `tags.json` | all tags | a tag cannot be deleted or moved, so a release tag keeps pointing at what was released | | `tag-creation.json` | all tags | only a repository admin can create a tag. Every tag starts the release, so a tag is a release to PyPI. A ruleset of its own, since the bypass of the admins must not extend to `tags.json` | Changing a policy means changing the json and applying it: ```bash .github/rulesets/apply.sh ``` The script is idempotent: it updates the rulesets which exist and creates the missing ones. It also sets the merge settings of the repository, i.e., auto-merge, delete branch on merge, and squash and rebase as the only merge methods. It needs the [github cli](https://cli.github.com) authenticated as a user with admin permission on the repository. ## Setup development environment Development needs [uv](https://docs.astral.sh/uv/) and a checkout of the repository: ```bash git clone https://github.com/matthiaskoenig/pkpdbib.git cd pkpdbib ``` A single sync creates the virtual environment in `.venv`, installs `pkpdbib` into it in editable mode and adds the complete tooling: ```bash uv sync --extra dev ``` The environment is resolved from `uv.lock`, which is committed, so local development, including `uv run ty check`, and the documentation build use the same versions. The tox environments, i.e. the test matrix and the `ty` check of continuous integration, resolve from `pyproject.toml` and do not use the lock; the lower bounds in `pyproject.toml` are what a user of the library installs against. After changing a dependency in `pyproject.toml` run `uv lock`: the `documentation` workflow syncs with `uv sync --locked`, which fails when `pyproject.toml` and the lock disagree. `uv lock --upgrade` moves the lock to the newest releases. The `dev` extra contains everything used below, i.e., pytest, ruff, ty, tox, pre-commit, zensical and bump-my-version, so nothing has to be installed separately. The python version is taken from `.python-version` (currently 3.14); to work against the oldest supported version instead use `uv sync --extra dev --python 3.13`, which replaces the environment. The tools are then run either with `uv run `, which uses the environment without activating it, or from the activated environment: ```bash source .venv/bin/activate # Linux and macOS .venv\Scripts\activate # Windows ``` The commands in this document are written without the `uv run` prefix; prepend it if the environment is not activated. The last step installs the git hook: ```bash uv run pre-commit install # install the hook, once per checkout uv run pre-commit run --all-files # check the current state of the repository ``` From now on every commit is checked with ruff (lint and format) and ty, i.e., the same checks that run in continuous integration. On a commit only the changed files are looked at, `--all-files` checks the whole repository and is what a newly added hook should be tried with. ## Testing The tests are written with pytest, tox runs them against every supported python version. They do not access the network: the download of scidownl and the Zotero client are replaced by fakes, so the suite is fast and deterministic. The tox environments are named after the interpreter (`py3.13` to `py3.15`, see `envlist` in `tox.ini`), a single one is run with ```bash tox r -e py3.15 ``` and the complete matrix, including the `ty` and `lowest` environments, in parallel with ```bash tox run-parallel ``` What follows `--` is passed to pytest, e.g., `tox r -e py3.14 -- tests/test_scihub_tools.py`. The environments need the interpreters, which uv installs with `uv python install 3.13 3.14 3.15`. Continuous integration runs the same environments with `uvx --with tox-uv tox -e py3.15`. The `lowest` environment installs the oldest version of every dependency which the lower bounds in `pyproject.toml` allow (`uv_resolution = lowest-direct`, their own dependencies stay at the newest) on python 3.13 and runs the suite against it, so a lower bound is only ever raised or lowered together with a run of it. It runs in the test matrix of `ci-cd.yml` and is part of the required `tests` check. ```bash tox r -e lowest ``` The workflows pin every action to the full commit SHA of a release, with the version in a comment; dependabot updates both. The tools run with `uvx` (tox, tox-uv, twine) are pinned in the `env` of `ci-cd.yml` and `ty.yml`, which dependabot does not update, so they are raised by hand. The same holds for uv itself, pinned with the `version` input of every `astral-sh/setup-uv` step in `ci-cd.yml`, `ty.yml` and `docs.yml`. `hatchling` in `[build-system]` stays a lower bound on purpose: the package is built in an isolated environment which resolves it anew, and the lower bound is the oldest release the build is known to work with. Dependabot also bumps the locked python dependencies (`uv.lock`, `versioning-strategy: lockfile-only` so the lower bounds in `pyproject.toml` are never raised by it) and the hook revisions in `.pre-commit-config.yaml`, each as one grouped weekly pull request. The `ruff` workflow installs the ruff version locked in `uv.lock`, so CI and `uv run ruff` agree, and the ruff and ty hook revisions are expected to match the lock; after merging one of these pull requests, raise the other to the same version if it lags. To run the tests directly against the development environment use ```bash pytest # the full suite pytest tests/test_scihub_tools.py # a single module pytest tests/test_scihub_tools.py::test_dois_from_json # a single test ``` ## Linting and formatting Linting and formatting use [ruff](https://docs.astral.sh/ruff/): ```bash ruff check # lint ruff format # format ``` ## Type checking Type checking is performed with [ty](https://docs.astral.sh/ty/): ```bash tox r -e ty ``` Or directly in the working tree: ```bash uvx ty check ``` The configuration lives in `[tool.ty]` in `pyproject.toml`. Warnings are treated as errors, so the codebase is kept free of diagnostics. Suppress an unavoidable diagnostic with a rule specific `# ty: ignore[rule-name]` rather than a blanket comment. ## Documentation The documentation is built with [Zensical](https://zensical.org/), the static site generator of the Material for MkDocs authors. The sources are markdown files in `docs/`, the site is configured in `zensical.toml` in the repository root. Nothing rendered is committed: the site is built by the `documentation` workflow on every push and published to [matthiaskoenig.github.io/pkpdbib](https://matthiaskoenig.github.io/pkpdbib) from the `develop` branch. Build the site into `site/`: ```bash uv run zensical build --clean ``` For writing, the preview rebuilds on save: ```bash uv run zensical serve ``` The API reference is rendered from the docstrings by [mkdocstrings](https://mkdocstrings.github.io/); a page in `docs/api/` only contains the module directive: ```markdown # scihub_tools ::: pkpdbib.scihub_tools ``` Docstrings are therefore the place to document functions and classes, the markdown files provide the narrative around them. Adding a module to the reference means adding such a page and an entry to `nav` in `zensical.toml`. ### Files for agents { #files-for-agents } Agents and language models read markdown, not rendered html. `scripts/llms_txt.py` writes the files of the [llms.txt convention](https://llmstxt.org/) into the built site, i.e., [llms.txt](https://matthiaskoenig.github.io/pkpdbib/llms.txt) as an annotated index of all pages, [llms-full.txt](https://matthiaskoenig.github.io/pkpdbib/llms-full.txt) with the complete documentation in a single file, and the markdown of every page next to its html (`/pdfs.md` for `/pdfs/`). The markdown of the API reference is generated from the docstrings with `inspect`, since the pages themselves only contain the mkdocstrings directive. ```bash uv run zensical build --clean uv run python scripts/llms_txt.py ``` The `documentation` workflow runs both steps, so the files are regenerated with every push. `docs/robots.txt` points crawlers at the sitemap and at these files. Zensical will provide agent context files itself at some point, then this script can go. ## Release A release is made from `develop`. Since `develop` only accepts pull requests, the release is prepared on a branch and tagged once that pull request is merged: 1. branch off `develop`: `git switch -c release/x.y.z develop` 2. write the release notes for the version in `release-notes/x.y.z.md` 3. make sure everything passes: `tox run-parallel`, `ruff check`, `tox r -e ty` 4. check the version bump: `uvx bump-my-version bump [major|minor|patch] --dry-run -vv` 5. bump the version: `uvx bump-my-version bump [major|minor|patch]`, which updates `src/pkpdbib/__init__.py` and `CITATION.cff` and commits. It does not create the tag; a squash or rebase merge would rewrite the commit and leave the tag behind on a commit which is not part of `develop` 6. push the branch, open the pull request against `develop` and merge it once the checks are green 7. tag the merged commit on `develop` and push the tag: ```bash git switch develop git pull git tag x.y.z git push origin x.y.z ``` Only a repository admin can push a tag, see `tag-creation.json`. This starts the `CI-CD` workflow, which runs the test matrix, builds the distributions in a job without write permissions (`build`), publishes them to [pypi](https://pypi.org/project/pkpdbib/) with attestations by trusted publishing from the `pypi` environment (`publish`), creates the GitHub release from `release-notes/x.y.z.md` (`github-release`) and fast-forwards `main` to the tagged commit (`sync-main`). Check the version before pushing, a tag cannot be moved or deleted afterwards. 8. test the installation from pypi in a fresh environment: ```bash uv venv --python 3.14 uv pip install pkpdbib ``` 9. once Zenodo has archived the release, update the citation information, i.e., `date-released` in `CITATION.cff` and the version, date and version DOI of the release in the citation of `README.md` and `docs/index.md`. `bump-my-version` only updates the version, not the date and the DOI, which are only known after the release. These changes go in through a pull request like everything else