← back to main

meiler lab / vanderbilt computational structural biology

a protein benchmark
that fails closed

may 2025 - present / remote from seattle / computational pipeline development + analysis

The BM5.5 project asks how Boltz-2 predictions change under matched physics-based relaxation, and whether those changes help complex geometry, interface quality, or only local stereochemistry. The hard part is not generating another figure. It is making every comparison traceable to the exact model, structure, protocol, score, and completion receipt that produced it.

Every lane is isolated until identity, method, and denominator checks pass. Archived Boltz-1 work remains available as a historical baseline but does not enter the new Boltz-2 claims.

A structure count is not a completion claim. The manuscript gate requires the exact Rosetta output set, all three postscore families, independent audits, and reproducible tables and figures.

01

benchmark and identity freeze

Each BM5.5 target, prediction lane, sample, relaxation method, replicate, and output path receives a collision-proof identity. Source hashes, method identifiers, environment captures, and expected denominators are fixed before analysis.

BM5.5SHA-256 manifestsidentity joins
02

Boltz-2 structure prediction

Boltz 2.2.1 runs with the Boltz-2 model, five diffusion samples, ten recycling steps, and 200 sampling steps. Two independently seeded lanes yield 1,285 models each across the 257 targets.

Boltz-22,570 modelsseeded lanes
03

matched AMBER relaxation

Every raw prediction is paired with the captured paper-method AlphaFold 2.3.2 AmberRelaxation implementation. The executable provenance, rather than remembered method wording, defines the canonical lane.

AlphaFold 2.3.2AmberRelaxationOpenMM
04

multi-protocol Rosetta relaxation

Six Rosetta protocol families and their replicate matrix are applied without pooling partial lanes. The complete planned release contains exactly 154,200 structure identities.

RosettaFastRelax154,200 outputs
05

three-axis structural scoring

TM-score measures global fold similarity, DockQ and its components resolve interface quality, and MolProbity measures local stereochemistry. Each completed structure must produce all three score rows.

TM-scoreDockQMolProbity
06

stratification and manuscript promotion

Only complete, validated tables feed comparisons by complex category, size, difficulty, prediction lane, and relaxation method. Main-text figures carry the shortest defensible result; exhaustive protocol detail remains in the supplement.

paired statisticsstratified analysisreproducible figures

The page reports fixed, audited milestones. It does not turn an active compute lane into a scientific conclusion.

raw Boltz-2 complete

Both prediction lanes are complete at 1,285 structures each, for 2,570 unique raw model identities. TM-score, DockQ, and MolProbity tables are complete at 7,710 metric rows.

complete / 2,570 structures + 7,710 metrics

paper-method AMBER complete

The matched relaxation lane is complete at 2,570 outputs, one for every raw Boltz-2 identity, with its own 7,710-row score release and immutable provenance.

complete / exact one-to-one pairing

Rosetta + postscore active

The canonical matrix is still progressing toward 154,200 Rosetta outputs and 462,600 corresponding metric rows. Partial counts remain operational status, not evidence for rankings.

active / final denominator not yet sealed

final comparisons withheld

No protocol winner or integrated Boltz-2 conclusion is promoted until the exact output and postscore receipts pass, followed by table, figure, reference, and manuscript checks.

fail closed / claims pending
why the gate matters

The June lab review emphasized that outliers cannot be silently rerun or discarded, interface quality must be decomposed, and results must be stratified by protein class and size. The current pipeline encodes those methodological requirements directly instead of repairing them after the figures are drawn.

historical benchmark lane, preserved

The earlier program generated more than 6,800 structures from AlphaFold 2.3.2 and Boltz-1 v0.4.1 across the same 257-complex BM5.5 benchmark. It evaluated six Rosetta 3.15 FastRelax protocols with five replicates each, followed by MolProbity and PoseBusters checks including clashscore, Ramachandran outliers, rotamer quality, Cβ deviations, bond lengths, bond angles, and steric clashes. Those results remain a documented baseline, but they are not presented as current Boltz-2 evidence.

global vs local A high TM-score can coexist with a poor protein-protein interface. DockQ components and MolProbity separate global fold preservation, binding-mode quality, and local stereochemical cleanup.
outlier behavior Targets that fail across every prediction method are analyzed differently from method-specific failures. Rerun criteria and exclusions are explicit so a blind user would know when a prediction is unreliable.
complex category Antibody-antigen, enzyme, and other complexes are stratified rather than averaged into one spread. The analysis asks where relaxation changes an interface, not only whether the pooled median moves.
protein size Small, medium, and large complexes are compared under the same identity and scoring rules to test whether apparent method effects are driven by system size.
protocol choice The final main-text comparison will use the smallest protocol view that supports the conclusion. The full six-protocol matrix and replicate behavior remain available in supplemental analyses.

The benchmark grew out of a broader Meiler Lab role supporting computational design in an ARPA-H-funded alphavirus vaccine program.

pipeline integration

ThermoMPNN, ESM, ProteinMPNN, MIF-ST, structure prediction, physics-based relaxation, and structural validation are connected into a reproducible path so the design team can move from a candidate sequence to a reviewable structure set for VEEV and MAYV envelope glycoproteins.

design-team support

The infrastructure is built for iteration: configurable batch submission, restart-safe outputs, explicit environments, and analysis tables that let experimental collaborators compare candidates without reconstructing the compute history.

my role

I built and maintain the prediction, relaxation, scoring, provenance, and figure-generation workflows; coordinate long-running HPC and workstation lanes; audit denominators and identities; and draft the first-author BM5.5 analysis under guidance from the Meiler Lab team.

in preparation

Benchmarking Boltz-2 Prediction and AMBER Relaxation on BM5.5 Protein Complexes

First author. Prediction, matched relaxation, interface scoring, and stratified analysis.

in preparation

alphavirus stabilization for vaccine design

Contributing author. Computational pipeline for the ARPA-H program.

presented aug 2025

VICB Summer Symposium talk

Alphavirus vaccine development through computational stabilization, Nashville.

watch the talk ↗

Protein_Data_Analysis

Structural scoring, MolProbity validation, geometry checks, comparative statistics, and figure generation.

view on github ↗

Protein_Relax_Pipeline

Configurable Rosetta and AMBER relaxation workflows with HPC batch submission and restart-safe outputs.

view on github ↗

Protein_Ideal

The earlier BM5.5 release and analysis scaffold. Its Boltz-1-era results remain an archived baseline rather than being relabeled as Boltz-2.

view on github ↗

prediction

Boltz-2AlphaFold 2.3.2ProteinMPNNThermoMPNNESMMIF-STBM5.5

relaxation

RosettaFastRelaxAMBEROpenMMAlphaFold AmberRelaxation

validation

TM-scoreDockQMolProbityPoseBustersPyMOL

analysis

PythonpandasNumPySciPyBioPythonpaired statisticsmatplotlibseabornJupyter

infrastructure

SLURMHPC arraysApple siliconmacOSBashLinuxGit

provenance

SHA-256 manifestsimmutable releasesidentity ledgerscompletion receiptsindependent audits
← back to research