REST API

Get sequences, organisms and evidence out of PlasticDB, and run an analysis against it — from a script.

Start here

Every endpoint lives under https://v2.plasticdb.org/api/v1 and answers JSON. Reading the database needs no account and no credential — every example on this page down to Run an analysis is a command you can paste into a terminal right now and get real data back. You only need a token to launch a job or submit data; that is further down, where it belongs.

Just want the data? Skip the API. Bulk downloads hand you the whole database — or any filtered slice of it — as FASTA, CSV, JSON or a ZIP, in one request.

From Python or R

Every curl on this page translates directly. The same request — wild-type enzymes with activity on PET — paging through the whole result:

# Python (pip install requests)
import requests

BASE = "https://v2.plasticdb.org/api/v1"
rows, offset = [], 0
while True:
    r = requests.get(f"{BASE}/proteins", params={
        "plastic": "PET", "kind": "wild_type", "limit": 500, "offset": offset,
    }, timeout=60)
    r.raise_for_status()          # a 422 names the bad parameter in r.json()["detail"]
    page = r.json()
    rows += page["items"]
    offset += len(page["items"])
    if not page["items"] or offset >= page["total"]:
        break
print(len(rows), "proteins")
# R (install.packages("httr2"))
library(httr2)

base <- "https://v2.plasticdb.org/api/v1"
rows <- list(); offset <- 0
repeat {
  page <- request(base) |>
    req_url_path_append("proteins") |>
    req_url_query(plastic = "PET", kind = "wild_type", limit = 500, offset = offset) |>
    req_perform() |>
    resp_body_json()
  rows <- c(rows, page$items)
  offset <- offset + length(page$items)
  if (length(page$items) == 0 || offset >= page$total) break
}
length(rows)

For a whole table, one bulk download is faster than paging: read.csv("https://v2.plasticdb.org/api/v1/downloads/proteins.csv?plastic=PET") or pandas.read_csv(…) on the same URL. To send a credential, add headers={"X-API-Key": key} (Python) or req_headers(`X-API-Key` = key) (R) — see Authentication.

What you probably came for

A complete, machine-readable listing of every route — including the curation and administrative ones this page does not cover — is generated from the running code: Swagger UI · ReDoc · openapi.json. This page is the curated subset a researcher actually needs.

Enzymes: records, sequences, structures

One enzyme, by accession

Every protein in PlasticDB has a stable accession — six characters beginning PB, such as PBG5RN. It is the identifier to quote in a paper and the one to key a script on.

curl https://v2.plasticdb.org/api/v1/proteins/PBG5RN

returns IsPETase from Pseudideonella sakaiensis:

{
  "id": 21,
  "accession": "PBG5RN",
  "name": "IsPETase",
  "kind": "wild_type",
  "uniprot_accession": "A0A0K8P6T7",
  "pdb_id": "6eqe",
  "ec_number": "3.1.1.101",
  "sequence": "MNFPRASRLMQAAVLGGLMAVSAAATAQTNPYARGPNPTA...",
  "source_microorganism": { "name": "Pseudideonella sakaiensis", "tax_id": 1547922, ... },
  "plastics": [ { "name": "PET", ... } ],
  "variants": [ ... 36 engineered versions ... ],
  "kinetics": [ ... ], "structures": [ ... ], "papers": [ ... ]
}
Never key on a protein name. Names are not unique — the database holds many rows called PETase, and two distinct records are both named IsPETase. accession is the unique handle; a name is a label for a human.

Accessions are case-insensitive: pbg5rn is PBG5RN. /proteins/{accession} answers 301 if that accession was merged into another record (follow the redirect), 301 to /variants/{accession} if it belongs to an engineered variant, 410 if it was withdrawn, and 404 if it never existed. /proteins/accession/{accession} is an equivalent alias that answers only {id, accession, name}. Most HTTP clients follow a 301 on their own; curl needs -L.

Its sequence, as FASTA

The FASTA and structure routes take the accession, like every other protein route. The numeric id from the record above still works too (/proteins/21/fasta).

curl https://v2.plasticdb.org/api/v1/proteins/PBG5RN/fasta
>PBG5RN | IsPETase | Pseudideonella sakaiensis | 2021 | PET
MNFPRASRLMQAAVLGGLMAVSAAATAQTNPYARGPNPTAASLEASAGPFTVRSFTVSRPSGYGAGTVYYPTNAGG...

Its structure, as PDB

curl -o PBG5RN.pdb https://v2.plasticdb.org/api/v1/proteins/PBG5RN/structure

Returns chemical/x-pdb when the record has a structure and 404 when it does not — check has_structure on the detail record first. Engineered variants have their own structures at /variants/{accession}/structure.

Many enzymes, filtered

Wild-type enzymes with published activity on PET, newest first:

curl "https://v2.plasticdb.org/api/v1/proteins?plastic=PET&kind=wild_type&limit=100"
{ "total": …, "items": [ { "id": ..., "accession": "PB...", "name": ..., "kind": "wild_type", ... } ] }

The paged list endpoints answer the same envelope: a total and an items array. (Counts change with every weekly release, so this page shows none; ask /stats for today's.) Some more filters worth knowing:

# free-text search (full-text over names, accession, gene and identifiers with kind=wild_type)
curl "https://v2.plasticdb.org/api/v1/proteins?q=PETase&kind=wild_type"

# engineered variants whose own kcat clears 5 s⁻¹
curl "https://v2.plasticdb.org/api/v1/proteins?kind=engineered&min_kcat=5"

# enzymes whose evidence rests on an HPLC assay
curl "https://v2.plasticdb.org/api/v1/proteins?evidence_method=hplc"

# everything derived from one organism
curl "https://v2.plasticdb.org/api/v1/proteins?microorganism=Pseudideonella"

The full filter set. How each one matches is stated per row, because it differs: an exact filter compares the whole value, a substring filter is a case-insensitive "contains" in which % and _ are ordinary characters, and a vocabulary filter takes one of a fixed list of values and answers 422 for anything else.

ParameterMeaning
qFree text. With kind=wild_type, a full-text search over name, accession, gene and identifiers, ranked by relevance, plus a substring match on the name. With kind=both or engineered, a substring of the name (a variant's own name).
nameSubstring of the name, case-insensitive: name=PETase also returns IsPETase and FAST-PETase.
accessionExact accession, case-insensitive. Matches a variant's own accession too.
plasticdb_idExact PlasticDB v1 identifier. A variant matches through its parent.
plastic · include_subtypesExact plastic name, abbreviation or full name, case-insensitive — PE or Polyethylene, never a substring, so PE does not match PET or LDPE. Add include_subtypes=true for its curated sub-types. The same rule on /microorganisms, /papers and every export.
microorganism · lineageSubstrings: of the source organism's name or any of its synonyms, and of its taxonomic lineage string (lineage=Pseudomonadota). On an engineered variant these resolve through its parent protein.
year · min_year · max_yearThe year the record was first reported (its earliest supporting paper): exactly year, or within the range.
substrate_form · evidence_methodVocabulary: the physical form of the substrate, and the assay the evidence rests on. Exact names from the vocabularies; an unknown value is a 422 listing the allowed ones. See below.
min_kcat · max_km · min_kcat_km · min_vmaxKinetic thresholds (kcat in s⁻¹, Km in mM, kcat/Km in M⁻¹·s⁻¹), matched against the record's own measurements: a variant on its own kinetics, a wild type on its own and never on its variants'. Stored units are normalised before comparing — a kcat reported in min⁻¹ is divided by 60 — and a measurement whose unit is not recognised is left out rather than compared as if it were s⁻¹. A threshold must be a finite number; NaN or inf is a 422.
min_tm · max_tmMelting temperature range (°C) from the stored enzyme properties. Both bounds apply to the same measurement; a variant is judged on its own Tm.
has_structuretrue: the record has its own experimentally determined structure with a PDB id (predicted AlphaFold/ESMFold models do not count). false: it has none.
has_sequencetrue / false: the record does / does not carry an amino-acid sequence.
ec_numberEC number prefix, matched on whole levels: 3.1.1 matches 3.1.1.101 and 3.1.1.-, never 3.1.10.x. A variant inherits its parent's EC number.
polarity · include_negativeVocabulary (positive, negative, inconclusive): records with at least one result of that polarity. Positive-only by default; include_negative=true drops the result filter and returns every public record.
kindwild_type, engineered or both (default). See below.
sort_by · sort_order · limit · offset · pageOrdering and paging. limit defaults to 100 and caps at 500. sort_by is one of year, name, accession, id (plus plasticdb_id, genbank_id with kind=wild_type); anything else is a 422. See Paging.

Engineered variants

PlasticDB treats an engineered variant as a first-class record with its own accession, its own kinetics and its own papers. To walk one enzyme's engineering lineage:

curl https://v2.plasticdb.org/api/v1/engineering/proteins/PBG5RN
{ "accession": "PBG5RN", "name": "IsPETase",
  "variants": [ { "id": 8, "accession": "PBM3M6", "mutations": [ { "raw": "N233K", ... } ], ... } ],
  "wild_type_kinetics": [ ... ], "wild_type_properties": [ ... ] }

One variant in full — mutations, engineering method, kinetics, structures, parent — by its accession or its numeric id (/variants/8 is the same record). Its own sequence and structure are at /variants/{accession}/fasta and /variants/{accession}/structure.

curl https://v2.plasticdb.org/api/v1/variants/PBM3M6

Or list every enzyme that has been engineered at all:

curl "https://v2.plasticdb.org/api/v1/engineering/proteins?limit=100"

Wild-type and engineered are one list, selected with kind

/proteins and the protein exports describe a union of two entity types:

kindReturns
both (default)Wild-type proteins followed by engineered variants.
wild_typeWild-type proteins only.
engineeredEngineered variants only.

Any other value is a 422, so a typo tells you rather than quietly returning the wrong dataset.

Every row carries a kind column. Because the two entity types have independent id sequences, the row key is (kind, id) — id alone is not unique in a kind=both result. Variant rows also carry parent_accession.

Every filter applies to both entity types. A variant resolves one of two ways. Inherited: microorganism, lineage and plasticdb_id have no column on a variant, so they resolve through its parent — ?microorganism=Pseudideonella&kind=engineered returns the Pseudideonella-derived variants. Its own: substrate_form, evidence_method, polarity and the kinetic thresholds match the variant's own experiments, never its parent's — ?min_kcat=5&kind=engineered selects variants whose own kcat clears 5 s⁻¹, which is the point of having engineered them.

Microorganisms

Organisms are addressed by their NCBI taxonomy id or by their PlasticDB accession (any case) — /microorganisms/1547922 and /microorganisms/PB3E8R are the same record:

curl https://v2.plasticdb.org/api/v1/microorganisms/1547922
{ "id": 2, "name": "Pseudideonella sakaiensis", "tax_id": 1547922, "accession": "PB3E8R",
  "degradation_links": [ { "plastic": { "name": "PET" }, "result": "positive",
                          "methods": [ "hplc", "sem", "enzyme_assay", ... ],
                          "positive_count": 8, "negative_count": 0,
                          "positive_papers": 3, "negative_papers": 0, "paper_count": 3 } ],
  "lineage": "...", "environments": [ ... ], "locations": [ ... ], "papers": [ ... ] }

*_count tallies interactions (one per strain or isolate a paper tested); *_papers counts the distinct papers behind them, which is the number the site's badges show. Expression hosts, knockouts and engineered strains (strain_occurrence.genetic_background) are never credited to the species; an organism record lists them under construct_interactions.

Everything reported to degrade PET:

curl "https://v2.plasticdb.org/api/v1/microorganisms?plastic=PET&limit=500"
MethodEndpointWhat it gives you
GET/microorganismsFilter by q, name, plastic, plastic_id, tax_id, temp_class, environment, location, year/min_year/max_year, lineage, the taxonomic ranks domain…genus, evidence_method, substrate_form, sole_carbon_source, aerobicity, polarity, include_negative, evidence_scope. Sort by name, tax_id or growth_temp_opt or id.
Exact: plastic (see above; include_subtypes=true widens it to sub-types), plastic_id, tax_id, year (first reported). Substring (case-insensitive): name (the name or any synonym), environment (the isolation-environment code, e.g. marine matches marine_sediment and marine_water), location (country or region), lineage and the rank filters (genus=Pseudomonas matches genus:Pseudomonas in the lineage). Vocabulary (422 for anything else): temp_class, evidence_method, substrate_form, sole_carbon_source, aerobicity, polarity, evidence_scope.
GET/microorganisms/{tax_id or accession}One organism: lineage, plastics with per-method evidence counts, proteins, papers. Takes evidence_scope (organism_or_enzyme or whole_organism).
GET/microorganisms/accession/{accession}The same record by PlasticDB accession, kept for existing links. An organism with no NCBI tax id is reachable by accession on either route.
GET/microorganisms/taxonomy/{rank}/{name}Everything at one point in the taxonomy, e.g. /taxonomy/genus/Bacillus. A genus or species that has been renamed and has no organism filed under it any more answers 302 to its current name (genus/Ideonella → genus/Pseudideonella); each top_microorganisms entry carries evidence_count, which is 0 for a member with no degradation result recorded.
GET/microorganisms/search?q=Type-ahead by name or tax id; up to 25 matches.
GET/microorganisms/namesEvery name and id, for populating a local lookup.
GET/microorganisms/lineage-levelsDistinct values per taxonomic rank, for building filter menus. Narrow with domain, kingdom, phylum, class, order.

evidence_scope decides whether an organism is credited with a plastic through its own assays only, or also through its enzymes. It defaults to the inclusive organism_or_enzyme; pass whole_organism for the strict reading, or enzyme_only for organisms credited solely through an enzyme. There is no value called organism. enzyme_only selects organisms, so it is accepted by /microorganisms and its exports only; the organism detail, /plastics, /plastics/{id}, /stats and /stats/charts take the first two and answer 422 for anything else.

Assay conditions. evidence_method (a method-vocabulary name such as co2_evolution), substrate_form (e.g. film, powder), sole_carbon_source (yes, no, not_stated) and aerobicity (aerobic, anaerobic, microaerophilic, not_reported — which also matches experiments with no value) must all hold in the same experiment, on an interaction that also satisfies plastic and polarity and is credited under evidence_scope. So ?plastic=LDPE&evidence_method=co2_evolution&sole_carbon_source=yes means a positive CO₂-evolution assay on LDPE in which LDPE was the sole carbon source.

temp_class (psychrophile <20 °C, mesophile 20–45 °C, thermophile ≥45 °C) buckets the organism's recorded optimum growth temperature (growth_temp_opt), not any assay temperature; organisms with no recorded optimum match none of the three. An unknown value (e.g. thermophilic) is a 422, as are unknown sole_carbon_source and aerobicity values.

Plastics, papers, compounds and pathways

MethodEndpointWhat it gives you
GET/plasticsEvery plastic in the vocabulary with its counts of degrading organisms and enzymes (distinct actors with a positive result). Filter with q (full-text), name, cas_number and polymer_class (substrings, case-insensitive — name=PET also returns PETG and PET blends; use plastic_id for one record), parent_id, and has_evidence (true: at least one degrader; false: none recorded). Sort by name, full_name, cas_number, microorganism_count or protein_count.
GET/plastics/{plastic_id}One polymer with its degraders and enzymes, and paper_count — the total /papers?plastic=<name> returns. Page the sub-lists with microorganism_offset/microorganism_limit and protein_offset/protein_limit (25 each by default, at most 500). microorganism_polarity/protein_polarity (positive, inconclusive, negative) pick which bucket is paged; microorganism_sort_by (papers, result, name) and protein_sort_by (also accession, plasticdb_id, id) order it. Unknown values are a 422.
GET/plastics/namesNames and ids only, for dropdowns.
GET/papersLiterature. Filter with q (full-text), year (exact) or min_year/max_year, doi, journal, authors and country (substrings, case-insensitive), and by the evidence a paper reports: plastic (exact name, abbreviation or full name — PE never matches PET; add include_subtypes=true for curated sub-types), microorganism (NCBI tax ID, accession or name; includes enzymes purified from it), protein (accession), evidence_method (a method name), result (positive, negative, inconclusive) and has_quantitative=true (at least one numeric observation). The evidence filters all apply to one reported result, so ?plastic=PE&evidence_method=weight_loss means a weight-loss assay on PE. An unknown evidence_method or result is a 422. limit defaults to 20 here, not 100.
GET/papers/{paper_id} · /papers/{paper_id}/footprintOne paper, and an aggregate of every record it contributed.
GET/compounds · /compounds/{compound_id}Monomers, additives and pathway intermediates. Filter with q and is_additive.
GET/pathways · /pathways/{pathway_id}Curated degradation pathways and their reaction steps.
GET/pathways/by-plastic/{plastic_id} · /by-plastic/{plastic_id}/graphThe pathways for one polymer, and the same thing shaped as a node/edge graph. [] means the plastic has no curated pathway; a plastic id that does not exist is a 404.
GET/stats · /stats/charts · /stats/by-countryDatabase totals, chart series and the country map.
GET/vocabulariesEvery controlled vocabulary in one response — methods, measurement types, environments, substrate forms, variant types, plastics, compounds. Cached an hour. See Controlled Vocabularies.
# PET, with its first 25 degraders and 25 enzymes
curl https://v2.plasticdb.org/api/v1/plastics/16

# the degradation pathways curated for PET
curl https://v2.plasticdb.org/api/v1/pathways/by-plastic/16

Read allowed values from /vocabularies at runtime rather than hard-coding them, so a client keeps working as the vocabularies grow.

Bulk downloads

The export endpoints under /api/v1/downloads are public and return the whole database or any filtered slice of it. They accept exactly the same filter parameters as the list endpoint they mirror, so an export always matches the result set you were looking at.

# every wild-type PET enzyme, as FASTA, ready for a DIAMOND database
curl -OJ "https://v2.plasticdb.org/api/v1/downloads/proteins.fasta?plastic=PET&kind=wild_type"

# the whole database in one archive
curl -OJ https://v2.plasticdb.org/api/v1/downloads/dataset.zip
EndpointContents
/downloads/dataset.zip · /dataset.jsonThe complete dataset, including protein_variants and the evidence spine — interactions, experiments and observations, keyed to each other and to the entity tables. The zip carries README.md (every table and how they join) and LICENSE.txt (CC BY 4.0); dataset.json opens with a metadata object carrying the same licence and citation. Takes no filters.
/downloads/proteins.fasta · /proteins.csv · /proteins.jsonEnzymes — wild-type and engineered, selected with kind. Accepts the whole /proteins filter set.
/downloads/microorganisms.csv · .jsonOrganisms. Accepts the whole /microorganisms filter set.
/downloads/plastics.csv · .jsonPolymers. Filters: q, name, cas_number, plastic_id, polymer_class, parent_id.
/downloads/papers.csv · .jsonLiterature (also served as /references.*). Accepts the /papers filters, including the evidence filters.
/downloads/protein-variants.csv · .jsonEngineered-variant detail: mutations, engineering method, kinetics, sequence (also served as /engineered-versions.*). Filters: parent_protein_id, parent_plasticdb_id, variant_type, engineering_method.
/downloads/search.zip · /search.jsonA search result set, exported. Requires q.

Joining a record to its sequence

Every row carries seq_id, which is exactly the identifier of that record's FASTA entry, and it is never empty. Join on seq_id — proteins.csv ↔ proteins.fasta ↔ protein-variants.csv all agree on it. accession is the real registry value and the better key when you specifically want registered records, but it is empty for a record whose accession has not been minted yet; seq_id falls back to PlasticDB_{id} (or PlasticDB_var_{id} for a variant) so those records stay joinable.

Variant rows populate their own interactions, kinetics, structures and papers; the source organism is inherited from the parent. Columns with no variant analogue (plasticdb_id, gene, genbank_id, uniprot_*, …) are empty.

Two deliberate differences. (1) proteins.fasta is a sequence file, so it omits records with no sequence; proteins.csv/.json keep them as metadata rows and can be one or more rows longer. (2) dataset.zip/dataset.json are a complete archive and apply no include_negative filter, so their proteins set is a superset of the endpoint's default — pass ?include_negative=true to match.

Selecting evidence by assay method

/proteins, /microorganisms and their exports accept evidence_method, matching the assay a record's evidence rests on — weight_loss, hplc, sem, clear_zone, ftir and so on, matched exactly against the method vocabulary.

curl -OJ "https://v2.plasticdb.org/api/v1/downloads/proteins.fasta?evidence_method=hplc"

This is a descriptive axis, not a quality ranking. Every record in PlasticDB is curated from published literature; no method ranks above another, and a record is not more or less trustworthy for having been assayed one way rather than another. Pick the assay that suits your analysis.

Checking a download before you pull it

HEAD works on every URL on this site, exports included: same status, same Content-Type and Content-Disposition, no body — so curl -I, link checkers and uptime monitors all work. No Content-Length is advertised, because exports are streamed and their size is not known until they are generated; a 0 there would be a lie rather than a number you could use. HEAD is deliberately absent from the OpenAPI schema: RFC 9110 defines it as GET minus the body, so listing a twin of every endpoint would double the reference without telling you anything new.

Which version of the data did you use?

PlasticDB publishes weekly. Curation runs continuously, but approved records stay invisible until a release makes them public — every Thursday at 06:20 UTC. The dataset version is the ISO year and week, such as 2026.39 (an extra out-of-cycle release in the same week gets a suffix, 2026.39.2), and it is deliberately not the application version in the page footer: 1.32.0 says which code is running, the dataset version says which data you searched.

curl https://v2.plasticdb.org/api/v1/releases/current
{
  "version": "2026.39",
  "released_at": "2026-09-24T06:23:11Z",
  "doi": null,
  "snapshot_release_label": null,
  "citation": "PlasticDB dataset release 2026.39. …",
  "license": "CC BY 4.0", ...
}

The shape, not live values: ask the endpoint for today's. doi and snapshot_release_label are null until a release has been archived with a DOI.

Record version in your methods section, and quote citation verbatim. That is the whole point of the field: it turns "downloaded from PlasticDB in August 2026" into a statement someone else can reproduce.

MethodEndpointWhat it gives you
GET/releases/currentThe live release: version, date, DOI, ready-made citation. Every field is null before the first release.
GET/releasesVersion history, newest first. limit defaults to 50 and caps at 200.
GET/releases/{version}One release and what it published — counts of promoted enzymes and organisms, timestamps.
GET/snapshots · /snapshots/latest · /snapshots/{id}The archived artefacts themselves: release_label, entity_counts, figshare_doi, checksum, file_size. /snapshots/latest is a 404 until a snapshot has been published.
A quiet week gets no new DOI. A DOI is minted only when the published data actually changed, so doi and snapshot_release_label may name an earlier release than version — that is correct, and it means the archive you cite genuinely contains the data you used. Full cadence and reasoning: docs/dataset-releases.md.

The exports under /downloads always serve the current release; there is no way to ask them for an older one. To pin an analysis to a past release, cite the Figshare DOI of that release's snapshot.

Run an analysis

Everything above is anonymous. This section is not: a job is owned by whoever launched it, so from here on you need a credential — one header, described in Authentication below. Get an API key once, then read on.

The fastest thing: search one sequence, synchronously

POST /bioinformatics/search-sequence runs DIAMOND against the PlasticDB protein database and returns the hits in the HTTP response. No job, no polling. Its parameters go in the query string, not in a JSON body.

curl -X POST -H "X-API-Key: $PLASTICDB_KEY" \
  "https://v2.plasticdb.org/api/v1/bioinformatics/search-sequence?sequence=MNFPRASRLMQAAVLGGLMAVSAAATAQ...&evalue=1e-5&pident=30"
{ "hits": [ { "subject_seq_id": "PBG5RN", "protein_name": "IsPETase", "entity_type": "protein",
              "pident": 99.7, "length": 290, "evalue": 1.2e-180, "bitscore": 601.2,
              "qstart": 1, "qend": 290, "sstart": 1, "send": 290 } ] }

A sequence long enough to overflow your proxy's URL limit will not fit in a query string. For anything larger than a single protein, upload it as a job instead.

Submitting a job

POST /jobs takes a multipart upload — the same path the web tools use:

curl -X POST https://v2.plasticdb.org/api/v1/jobs \
  -H "X-API-Key: $PLASTICDB_KEY" \
  -F "job_type=annotate_genome" \
  -F "file=@my_genome.faa" \
  -F "evalue=1e-5" -F "pident=30" -F "blast_type=blastp"
{ "id": "5f2c…", "job_type": "annotate_genome", "status": "waiting", "result_count": 0, ... }
job_typeWhat it doesUpload
annotate_geneAlign one or a few sequences against PlasticDB.file (FASTA)
annotate_genomeScreen a whole proteome or genome for plastic-active enzymes.file (FASTA)
annotate_listAnnotate a TSV list of organism names.file (TSV) plus genus_column, species_column, field_separator
compare_genomesCompare the plastic-degrading repertoire of several genomes.files — two or more, repeat the field
pathway_analysisMap hits onto curated degradation pathways.file (FASTA)
hmm_screenSearch profile HMMs built from families of curated plastic-active enzymes — reaches remote homologs a pairwise alignment misses.file (FASTA, protein) plus optional target_polymers

One pipeline is not reached this way. The Taxonomy Tree is POST /tree/build, which takes a microorganism filter or a list of ids rather than a sequence upload, and needs no credential. It is deliberately not a job_type here: this endpoint saves whatever you upload as the job's input FASTA, and the tree builder never reads it, so a build_tree submission used to return 201 and then build something unrelated to what was sent (#459).

target_polymers applies to hmm_screen only — repeat the field to send several, and omit it to summarise every polymer. It narrows the per-polymer tally on the job; every hit is stored either way. An unrecognised name is a 422 listing the accepted set. A hit is a profile match, not a measurement: it says a sequence resembles a family whose members have been reported acting on a polymer, not that it degrades one.

Parameters, all optional: evalue (default 1e-5, must be in the range 0 < e ≤ 10), pident (default 30, 0–100), blast_type (blastp or blastx), organism_type (no_signalp, arch, gram+, gram-, euk) and include_variants (default false — off so results are natural degraders rather than a protein's own mutants; it is recorded on the job so a result set can be read months later without guessing which database produced it).

Getting the results back

curl -H "X-API-Key: $PLASTICDB_KEY" https://v2.plasticdb.org/api/v1/jobs/5f2c…
curl -OJ -H "X-API-Key: $PLASTICDB_KEY" https://v2.plasticdb.org/api/v1/jobs/5f2c…/download
MethodEndpointWhat it gives you
POST/jobsSubmit a job. 30/hour.
GET/jobsYour jobs, newest first.
GET/jobs/{job_id}Status and results, including the DIAMOND hits.
GET/jobs/{job_id}/downloadThe results file — a TSV, or for compare_genomes a ZIP of comparison_matrix.tsv plus one hits TSV per genome.
GET/jobs/{job_id}/comparison?format=tsv|csvcompare_genomes only: the genome × plastic matrix as a file (404 for any other job).

Each hit's protein carries the matched record's source_microorganism (null when none is on record or it is not yet public), its own degradation_links per plastic with result counts (methods is left empty here — read it from /proteins/{accession}), and paper_count. The paper list itself is not embedded per hit; /proteins/{accession} has it.

A finished compare_genomes job also returns comparison: the plastics (columns) and one row per genome with counts (query sequences with a hit to a record that has a positive result for that plastic, each sequence once per plastic), normalised (per 1,000 input proteins for blastp, per megabase for blastx; null for jobs run before genome sizes were recorded), positive, and not_positive — sequences whose hits carry no positive result for any plastic, which are never counted as potential. It is built from the hits and their current evidence each time it is read.

status runs waiting → running → done, or error with an error_message. Poll the detail endpoint. Two fields tell you where the output is: result_count counts the structured hits in the response, and has_result_file says whether /download will return anything — annotate_list and pathway_analysis write a file instead of structured rows, so their result_count is 0 even when the job found plenty.

A job you launched is yours: GET /jobs/{job_id} and its /download answer 403 to anyone else. Jobs launched anonymously from the web tools have no owner and are readable by anyone holding the id, which is why the read routes accept a request with no credential at all.

The JSON job endpoints

/bioinformatics/* mirrors the same pipelines without a multipart upload, taking base64 FASTA in the query string — built for scripted and agent access. All require a credential and are limited to 30/hour.

EndpointRequired parameter
POST /bioinformatics/annotate-genefasta_b64 — the FASTA file, base64-encoded. Plus the same evalue, pident, blast_type, organism_type, include_variants as above.
POST /bioinformatics/annotate-genome
POST /bioinformatics/annotate-list
POST /bioinformatics/pathway-analysis
POST /bioinformatics/compare-genomesfasta_b64 repeated, once per genome (at least two), plus optional labels repeated in the same order to name them (default genome_1, genome_2, …). Names must be distinct.
POST /bioinformatics/hmm-screenfasta_b64, plus target_polymers (repeat the parameter; omit to screen every profile).
POST /bioinformatics/search-sequencesequence — a plain sequence string. Synchronous; returns hits, not a job.

Each returns the same job object as POST /jobs, so poll /jobs/{job_id} the same way — except /search-sequence, which answers with the hits themselves. The Taxonomy Tree builder is separate: POST /tree/build takes a microorganism filter or an explicit id list, works without a credential, and the result is read from GET /tree/{job_id}: an NCBI-taxonomy cladogram (ranks only, no branch lengths) in result_newick, with the per-organism plastic matrix and the counts of organisms left out in result_metadata.

Authentication

You need this for exactly three things: launching jobs, submitting data, and your own account. Nothing above Run an analysis requires it.

Get an API key

Register, confirm your email, then generate a key at Profile › API Keys. It is shown once and never expires until you revoke it, which is what makes it the right credential for a script or a pipeline. Send it as a header:

curl -H "X-API-Key: $PLASTICDB_KEY" https://v2.plasticdb.org/api/v1/jobs

Or exchange your password for a short-lived JWT and send that as a bearer token:

curl -X POST https://v2.plasticdb.org/api/v1/auth/login \
  -d "username=you@example.org" -d "password=…"
# → { "access_token": "eyJhbG…", "token_type": "bearer" }

curl -H "Authorization: Bearer eyJhbG…" https://v2.plasticdb.org/api/v1/auth/me

An API key is also accepted as a bearer token, so a single Authorization: Bearer code path works for both if you prefer one.

Being logged in to the website does not authenticate an API call

The browser session is an httpOnly cookie, and the API deliberately refuses it — a cookie is attached by the browser automatically, so accepting one would make every API route reachable by a hostile page on another site (CSRF). The pages you click through use the cookie; the API uses a header you have to set on purpose. Practically: a fetch() from a page must send Authorization: Bearer …, and a plain <a download> link cannot authenticate at all, because a link cannot set a header.

MethodEndpointAuthDescription
POST/auth/registerPublicCreate an account, inactive until email confirmation. Body: {email, password, full_name}. 5/hour.
POST/auth/loginPublicGet a JWT. Form body: {username: email, password}. Returns {access_token, token_type}. 10/minute.
GET/auth/confirm-email?token=PublicConfirm a registration email.
POST/auth/resend-confirmationPublicResend the confirmation email. 3/hour.
POST/auth/forgot-password · /auth/reset-passwordPublicPassword-reset request and completion. 5/hour each.
POST/auth/logoutPublicClear the browser session cookie. Irrelevant to a header-authenticated client.
GET · PATCH/auth/meAuthRead or update your profile. PATCH body: {full_name, password}.
POST · DELETE/auth/api-keyAuthGenerate a new API key — returned once, save it — or revoke the current one.

Contribute data

Submissions are curated before they appear, and are published by the next weekly release. The record shape is documented at Contributing Data.

MethodEndpointAuthDescription
POST/submissionsAuthSubmit interactions with their actors, experiments and observations. Requires "license_accepted": true: accepted records are published under CC BY 4.0. 30/hour.
GET/submissions/mineAuthYour submissions and their curation status.
GET/submissions/{id}/matchesAuthExisting records your submission may duplicate.
POST/submissions/lookup-doiPublicPaper metadata by DOI — the database first, then Crossref. Body: {doi}. 30/hour.
GET/submissions/lookup-taxid?name=PublicSearch NCBI taxonomy by name. 30/hour.
POST/submissions/lookup-sequenceAuthFetch a sequence by accession. Body: {accession_id, database: genbank|uniprot}. 30/minute.
POST/submissions/validate-sequenceAuthCheck whether a string is a valid amino-acid or nucleotide sequence. 30/minute.
GET/submissions/isolation-environments · /sample-types · /evidence-types · /methodsPublicVocabulary helpers for building a submission.

Reference

Paging and identifiers

List endpoints return { "total": …, "items": [ … ] } and page with limit and offset. limit caps at 500 almost everywhere — the exceptions are /releases and /snapshots (200) and /search/grouped (100). Defaults differ: 100 for /proteins, /microorganisms, /plastics, /compounds and /pathways; 50 for /jobs, /releases and /snapshots; 20 for /papers and /search; 10 for /search/grouped.

Paging is stable. Every sort ends with a unique key (the record id), so walking offset=0, limit, 2×limit, … returns each record exactly once however many records share a year or a name. page (1-based) is accepted as shorthand for offset=(page−1)×limit; send page or a non-zero offset, not both (422). sort_by and sort_order take a fixed set of values, listed per endpoint in the OpenAPI schema; an unknown one is a 422 naming the allowed values, never a silent fallback to the default order.

Sorting. Text sorts ignore case (uncultured bacterium sorts among the Us, not after Z), and records with no value for the sort key come last in both directions — so sort_by=growth_temp_opt&sort_order=desc opens with the hottest organism, not with the hundreds that have no recorded temperature.

Identifiers. Dates are ISO 8601. Proteins, engineered variants and organisms all carry a PlasticDB accession (PB + 4 characters), and it works on every route for its record, in any case: /proteins/{accession} (and /fasta, /structure), /variants/{accession}, /microorganisms/{accession}. Those routes also still take the numeric id (an organism's NCBI tax_id). Plastics, papers, compounds and pathways have integer ids only; jobs and users are UUIDs.

Unknown query parameters are ignored, not rejected. A misspelled parameter name (?plastc=PET) is dropped and the request answers as if it had not been sent — a misspelled value of a known parameter is a 422. This is deliberate: rejecting unknown names would break every client that sends a parameter a later release renames or retires. If a filter seems to have no effect, check its spelling against the OpenAPI schema.

Rate limits

Limits are set per route, and only on the routes listed below; every other route — the detail routes for plastics, microorganisms, proteins, papers and compounds, /stats, /vocabularies, /releases, /snapshots — is not rate-limited at all. Each listed route keeps its own count, so 200 requests to /proteins do not use up /papers.

Who shares a count. A request carrying a login token (Authorization: Bearer with the JWT from /auth/login) is counted for that user. Every other request — anonymous, or authenticated with an API key (in X-API-Key or as a bearer token) — is counted for its IP address, so everyone behind one institutional gateway shares it. Over the limit is a 429 with {"detail": "Rate limit exceeded: …"}; wait for the window to roll over and retry. No rate-limit headers are sent, so pace a script yourself — or use the bulk downloads, which return a whole filtered set in one request.

Endpoint(s)Limit
/proteins, /microorganisms, /plastics, /papers, /compounds (the lists, not the detail routes), /search, /search/grouped, every /downloads/* export, /variants/*, /engineering/*, /pathways/by-plastic/*/graph200 / hour
/search/autocomplete300 / hour
POST /jobs, all POST /bioinformatics/*, POST /tree/build, POST /submissions, POST /annotations, POST /submissions/lookup-doi, GET /submissions/lookup-taxid30 / hour
POST /submissions/lookup-sequence, /validate-sequence30 / minute
POST /auth/login10 / minute
POST /auth/register, /forgot-password, /reset-password5 / hour
POST /auth/resend-confirmation3 / hour

Status codes

CodeMeans
301This accession was merged into another record, or it belongs to an engineered variant and the record is at /variants/{accession}. Follow the Location header.
302A taxonomy name that is now a synonym (see /microorganisms/taxonomy/{rank}/{name}). Location is the current name. Temporary on purpose: the mapping follows the data.
401No credential, or one the server does not recognise. See Authentication — a browser session cookie is not a credential here.
403Authenticated, but the resource is someone else's — most often another user's job.
404No such record, or the record is not public.
410This accession was withdrawn. It will not come back, and it has no successor.
422A parameter value is out of range or misspelled — ?kind=enginered, ?evidence_scope=organism, evalue=0. Deliberately an error rather than a silent fallback to the default, which would hand you the wrong dataset with a 200.
429Rate limited.

Error bodies

Every error is a JSON object with a detail key. On a 422 from a read (GET) endpoint, detail is always a list, one entry per problem, each naming the parameter in loc:

{ "detail": [ { "type": "enum", "loc": ["query", "evidence_scope"],
                "msg": "Input should be 'organism_or_enzyme' or 'whole_organism'",
                "input": "organism" } ] }

Every other status carries a human-readable string: {"detail": "Protein not found"}. The POST job and submission routes answer some of their own checks (an empty upload, a bad e-value) with a string detail too; read detail as "a list of problems, or one sentence" and you will handle both.

Not covered here

The curation, ingestion, annotation, forum, chat and administrative routes are site machinery rather than a research interface, and most require a curator or superuser account. They are all in the generated schema if you need them.