Reproducibility¶
Most papers in structural biology don't ship reproducible code. molforge
already records what produced an output — every engine wrapper attaches a
Provenance (engine, version, parameters, inputs,
and a pointer to the step it consumed) to result.metadata["provenance"].
molforge.reproducibility turns that
chain into a single, human-readable pipeline.yaml — the artifact a
methods section can point at.
from molforge.reproducibility import emit_pipeline
folded = esmfold.predict(sequence)
docked = vina.dock(folded, ligand)
emit_pipeline(docked, "pipeline.yaml")
The artifact¶
The manifest linearizes the provenance chain into ordered steps and adds a consolidated environment block:
molforge_pipeline: 1
generated: "2026-07-15T12:00:00+00:00"
environment:
molforge_version: "0.8.0"
python_version: "3.12.13"
platform: "macOS-14.3-arm64"
engines: {ESMFold: "1.0.3", Vina: "1.2.5"}
steps:
- step: 1
engine: ESMFold
engine_version: "1.0.3"
inputs: {sequence: "MKT..."}
parameters: {recycles: 4}
- step: 2
engine: Vina
engine_version: "1.2.5"
inputs: {ligand: "lig.sdf"}
parameters: {exhaustiveness: 8}
output: {type: DockingResult}
Inspecting a manifest¶
You don't have to write a file to use it. Build the manifest in memory and inspect it:
from molforge.reproducibility import pipeline_manifest
m = pipeline_manifest(docked)
print(m.describe())
# pipeline (2 steps) — molforge 0.8.0
# 1. ESMFold v1.0.3
# 2. Vina v1.2.5
m.steps # list[PipelineStep]
m.environment # the environment block
m.to_dict() # plain dict — for logging, comparison, a DataFrame
pipeline_manifest (and emit_pipeline) accept any molforge output that
carries provenance — a Protein, DockingResult, Pose,
DesignedSequence, ... — or a Provenance instance directly.
Engine and backend versions¶
A step records the version its wrapper could read at the time. For a
pip-installed engine that is exact. For a wrapper that shells out to a native
binary — fpocket, P2Rank, GROMACS, sander, gnina — there is nothing to read,
so those steps carry no version and the manifest has a blank exactly where a
reader wants a number.
engine_versions is the registry that answers the question directly:
from molforge import engine_versions
for backend in engine_versions().values():
if backend.available:
print(f"{backend.name:14s} {backend.version or '(no version flag)'}")
Every backend molforge can drive is reported, installed or not — an absent
engine is a fact about the environment worth recording. Each
BackendVersion carries its category (folding, docking, pockets,
md, freeenergy, generative, runtime), how it was found (kind:
python, executable, or repo), whether it is available, the resolved
location for a binary, and a detail explaining any empty version. Those
are three different reasons a version can be missing and the registry keeps
them apart:
kind |
Found by | Version from |
|---|---|---|
python |
importlib.metadata |
the distribution — exact, free |
executable |
shutil.which |
running its version flag, where it has one |
repo |
nothing — the caller passes repo_dir |
undetectable; said so explicitly |
Nothing here raises, on any machine, with none of it installed.
The manifest uses the registry to fill its own blanks, under a separate
engines_detected key:
environment:
engines: {ESMFold: "1.0.3"} # what ran, recorded at the time
engines_detected: {fpocket: "4.1"} # what is installed now
They are kept apart because engines_detected is the weaker claim: it
describes the machine when the manifest was written, which is not necessarily
what produced the output. If the run is long and the record matters, capture
engine_versions() up front rather than relying on the backfill.
Probing a native binary means running it, so the sweep spawns subprocesses —
for installed tools only, with a timeout, and memoized. Pass
probe_executables=False to skip every subprocess and still get availability
from $PATH.
Formats and the repro extra¶
The in-memory manifest and its to_dict() / to_json() forms need no
third-party dependency — they're part of molforge's numpy-only core:
Reading and writing the .yaml form needs PyYAML, an opt-in extra so
the core stays light:
Without it, to_yaml() / a .yaml load_pipeline raise a clear
ImportError with that install hint; JSON keeps working regardless. Load a
manifest back with load_pipeline, which
picks the format by file suffix.
Replaying a pipeline¶
replay() re-executes a manifest's chain, threading each step's output into
the next:
from molforge.reproducibility import load_pipeline, replay
manifest = load_pipeline("pipeline.yaml")
result = replay(manifest, context={"ligand": "aspirin.sdf"})
It resolves each step's engine from the registry (molforge's own wrappers,
plus anything registered under molforge.plugins),
reconstructs the call from the recorded parameters, and runs it. Each
operation (predict, dock, …) has a replay handler that owns its
reconstruction, so the "which recorded input is the previous step's output
vs. a literal" wiring is handled per operation rather than guessed — a
docking step's receptor is threaded from the fold that preceded it, its
ligand comes from context or the recorded value.
Register a handler for a custom operation with register_replay_handler:
from molforge.reproducibility import register_replay_handler
@register_replay_handler("my_op")
def _replay_my_op(engine_factory, step, upstream_output, context):
...
Replay is inherently partial: the engines must be installed, GPU steps need
the hardware (replay orchestrates, it doesn't provide compute), and inputs
that aren't literals must be supplied via context. A missing engine, an
operation with no handler, or an unresolvable input raises a clear
ReplayError. molforge ships predict and dock handlers in v1.
What v1 doesn't do¶
- Single output, linear chain. A manifest describes one output's
provenance chain (provenance has a single parent pointer). Merging
several outputs' chains — e.g. an entire
DesignTable— into one manifest with shared-ancestry deduplication is a future extension.