Skip to content
OpenAlgo

OpenAlgo ResearchPrivate beta

Stop debugging PDFs. Start doing science.

Paste a DOI or upload a PDF. OpenAlgo extracts the method, makes you check every value against its source, and generates a clean Python project from tested templates, with a report of everything the authors left out.

Private beta: approved accounts translate today · Request access

REPRODUCIBILITY_GAPS.md

5 specified5 missing

Specified in paper

  • Dataset

    PubChem, ChEMBL, eMolecules

  • Molecular features

    RDKit molecular descriptors

  • Model

    Random forest, XGBoost

  • Split strategy

    80/20, stratified

  • Primary metric

    R²

Not specified — flagged

  • Random seed

    Not mentioned — default 42

  • Salt stripping

    Not discussed — RDKit default

  • Stereoisomers

    Not addressed — kept as-is

  • Hyperparameter search

    Not described — template defaults

  • RDKit version

    Not specified — latest stable

Fig. 1 — Shipped in every generated project.

papers with public gap reports
100papers with public gap reports
generated repository contracts
10generated repository contracts
template families
5template families
independent reads per paper
2independent reads per paper

Why this works

QSAR pipelines are standardized. That’s the point.

Almost every QSAR paper follows one architecture: molecules in, featurization, a split, a model, metrics out. Because the structure rarely changes, mapping paper text to working Python is a constrained problem, so it can be solved reliably instead of improvised.

Input01

SMILES + targets

Dataset, source, cleaning

Featurize02

Fingerprints · descriptors · graphs

Morgan radius, bits, RDKit set

Split03

Random · scaffold · temporal

Ratios, stratification, seed

Model04

RF · XGBoost · SVM · GNN

Architecture, hyperparameters

Evaluate05

RMSE · R² · ROC-AUC

Matched to what the paper reports

Workflow

Extract. Review. Generate.

AI does the reading. You make the calls. Templates write the code.

Provenance

field: split_strategy

Extracted value

Scaffold split · 80 / 10 / 10

“…molecules were partitioned with a Bemis–Murcko scaffold split (80/10/10) to limit leakage between train and test…”Section 2.3 · page 4

Read A

✓ scaffold 80/10/10

Read B

✓ scaffold 80/10/10

Confidence: high — both reads agree, quoted source found

In motion

Minutes from DOI to a project you can defend.

The tedious part — environments, loaders, split logic, config — is done for you, so review time goes to the decisions that need a scientist.

Fig. 1 — One paper, translated

  1. extract
  2. review
  3. generate
Illustrative. Real sessions link every field to the sentence it came from.Templates, not freeform generation

Not another chatbot

A compiler for research papers, not a code guesser.

General-purpose models are impressive and unaccountable. OpenAlgo uses AI where it is strong — reading — and deterministic templates where you need guarantees.

 General-purpose chatbotOpenAlgo Research
Where the code comes fromWritten from scratch on every run✓Tested templates filled with values you reviewed
ProvenanceNone✓Every value linked to the sentence it came from
When it misreadsSilently wrong✓Two independent reads; disagreements shown to you
When the paper is silentGuesses, and hides the guess✓Flags the gap and names the default used
Chemistry defaultsGeneric✓RDKit setup, salt stripping, scaffold splits built in
Run it twiceDifferent code each time✓Same reviewed paper, same project

Coverage today

Five template families

  • Fingerprint / descriptor classificationStable
  • Fingerprint / descriptor regressionStable
  • Graph neural network classificationBeta
  • Graph neural network regressionBeta
  • Evaluation / split wrapperStable
Limits of each family

Reproducibility Hub

Proof, in public.

Before asking anyone to trust us, we published gap reports for 100 recent QSAR and molecular ML papers, with citation snapshots, template fit and the details each paper omits. Read them before you translate your own.

Transparency

What we can’t do — yet.

  • 01Papers omit details: random seeds, exact hyperparameter grids, manual cleaning steps. We can’t extract what the authors never wrote down, but we flag what’s missing so you know where to look.
  • 02If the dataset isn’t published, you get the full pipeline with placeholder data loading, clearly marked, ready for your own files.
  • 03Novel architectures that don’t exist in standard libraries are outside the templates. When a paper needs custom math, we tell you up front.
  • 04Docking, molecular dynamics, genomics and retrosynthesis are out of scope for now.

Questions

Before you trust it

Chatbots generate code from scratch every time: different structure, different bugs, no chemistry-specific defaults. OpenAlgo uses tested templates with built-in RDKit configuration, salt stripping, scaffold splits and descriptor scaling. The code is consistent and auditable, not a one-off you have to debug.

Every extracted value is surfaced for review before any code is generated. Two independent reads with different strategies flag any field where they disagree; you see both values and pick. Nothing reaches generated code without your sign-off.

Every paper gets a source-quality scorecard: dataset size, whether the data is public, whether external validation was reported, and red flags such as very high accuracy on noisy benchmarks or small datasets paired with complex models. It is informational. You are the scientist.

You get the full pipeline — featurization, splitting, training, evaluation — with placeholder data loading. Data format, column names and file paths are clearly marked so you can plug in your own files and run it.

We onboard teams with real papers to evaluate, so we can hold the output to a scientific standard. Translation is open to approved accounts; the Reproducibility Hub is public to everyone. Request access to join the next cohort.

Your next paper is already waiting.

Find out in minutes whether it holds up, and walk away with a project you can run, review and extend.