REST API
Get sequences, organisms and evidence out of PlasticDB, and run an analysis against it — from a script.
Start here
Every endpoint lives under https://v2.plasticdb.org/api/v1 and answers JSON.
Reading the database needs no account and no credential — every example
on this page down to Run an analysis is a command you can paste into a
terminal right now and get real data back. You only need a token to launch a job or submit
data; that is further down, where it belongs.
From Python or R
Every curl on this page translates directly. The same request — wild-type
enzymes with activity on PET — paging through the whole result:
# Python (pip install requests)
import requests
BASE = "https://v2.plasticdb.org/api/v1"
rows, offset = [], 0
while True:
r = requests.get(f"{BASE}/proteins", params={
"plastic": "PET", "kind": "wild_type", "limit": 500, "offset": offset,
}, timeout=60)
r.raise_for_status() # a 422 names the bad parameter in r.json()["detail"]
page = r.json()
rows += page["items"]
offset += len(page["items"])
if not page["items"] or offset >= page["total"]:
break
print(len(rows), "proteins")
# R (install.packages("httr2"))
library(httr2)
base <- "https://v2.plasticdb.org/api/v1"
rows <- list(); offset <- 0
repeat {
page <- request(base) |>
req_url_path_append("proteins") |>
req_url_query(plastic = "PET", kind = "wild_type", limit = 500, offset = offset) |>
req_perform() |>
resp_body_json()
rows <- c(rows, page$items)
offset <- offset + length(page$items)
if (length(page$items) == 0 || offset >= page$total) break
}
length(rows)
For a whole table, one bulk download is faster than paging:
read.csv("https://v2.plasticdb.org/api/v1/downloads/proteins.csv?plastic=PET")
or pandas.read_csv(…) on the same URL. To send a credential, add
headers={"X-API-Key": key} (Python) or
req_headers(`X-API-Key` = key) (R) — see Authentication.
What you probably came for
A complete, machine-readable listing of every route — including the curation and administrative ones this page does not cover — is generated from the running code: Swagger UI · ReDoc · openapi.json. This page is the curated subset a researcher actually needs.
Enzymes: records, sequences, structures
One enzyme, by accession
Every protein in PlasticDB has a stable accession — six characters beginning
PB, such as PBG5RN. It is the identifier to quote in a paper and
the one to key a script on.
curl https://v2.plasticdb.org/api/v1/proteins/PBG5RN
returns IsPETase from Pseudideonella sakaiensis:
{
"id": 21,
"accession": "PBG5RN",
"name": "IsPETase",
"kind": "wild_type",
"uniprot_accession": "A0A0K8P6T7",
"pdb_id": "6eqe",
"ec_number": "3.1.1.101",
"sequence": "MNFPRASRLMQAAVLGGLMAVSAAATAQTNPYARGPNPTA...",
"source_microorganism": { "name": "Pseudideonella sakaiensis", "tax_id": 1547922, ... },
"plastics": [ { "name": "PET", ... } ],
"variants": [ ... 36 engineered versions ... ],
"kinetics": [ ... ], "structures": [ ... ], "papers": [ ... ]
}
PETase, and two distinct records are both named
IsPETase. accession is the unique handle; a name is a label for
a human.
Accessions are case-insensitive: pbg5rn is PBG5RN.
/proteins/{accession} answers 301 if that accession was merged
into another record (follow the redirect), 301 to
/variants/{accession} if it belongs to an engineered
variant, 410 if it was withdrawn, and 404 if it never
existed. /proteins/accession/{accession} is an equivalent alias that answers
only {id, accession, name}. Most HTTP clients follow a 301 on
their own; curl needs -L.
Its sequence, as FASTA
The FASTA and structure routes take the accession, like every other protein route. The
numeric id from the record above still works too
(/proteins/21/fasta).
curl https://v2.plasticdb.org/api/v1/proteins/PBG5RN/fasta
>PBG5RN | IsPETase | Pseudideonella sakaiensis | 2021 | PET MNFPRASRLMQAAVLGGLMAVSAAATAQTNPYARGPNPTAASLEASAGPFTVRSFTVSRPSGYGAGTVYYPTNAGG...
Its structure, as PDB
curl -o PBG5RN.pdb https://v2.plasticdb.org/api/v1/proteins/PBG5RN/structure
Returns chemical/x-pdb when the record has a structure and 404
when it does not — check has_structure on the detail record first. Engineered
variants have their own structures at /variants/{accession}/structure.
Many enzymes, filtered
Wild-type enzymes with published activity on PET, newest first:
curl "https://v2.plasticdb.org/api/v1/proteins?plastic=PET&kind=wild_type&limit=100"
{ "total": …, "items": [ { "id": ..., "accession": "PB...", "name": ..., "kind": "wild_type", ... } ] }
The paged list endpoints answer the same envelope: a total and an
items array. (Counts change with every weekly release, so this page shows
none; ask /stats for today's.) Some more filters worth knowing:
# free-text search (full-text over names, accession, gene and identifiers with kind=wild_type) curl "https://v2.plasticdb.org/api/v1/proteins?q=PETase&kind=wild_type" # engineered variants whose own kcat clears 5 s⁻¹ curl "https://v2.plasticdb.org/api/v1/proteins?kind=engineered&min_kcat=5" # enzymes whose evidence rests on an HPLC assay curl "https://v2.plasticdb.org/api/v1/proteins?evidence_method=hplc" # everything derived from one organism curl "https://v2.plasticdb.org/api/v1/proteins?microorganism=Pseudideonella"
The full filter set. How each one matches is stated per row, because it
differs: an exact filter compares the whole value, a substring filter is
a case-insensitive "contains" in which % and _ are ordinary
characters, and a vocabulary filter takes one of a fixed list of values and
answers 422 for anything else.
| Parameter | Meaning |
|---|---|
q | Free text. With kind=wild_type, a full-text search over name, accession, gene and identifiers, ranked by relevance, plus a substring match on the name. With kind=both or engineered, a substring of the name (a variant's own name). |
name | Substring of the name, case-insensitive: name=PETase also returns IsPETase and FAST-PETase. |
accession | Exact accession, case-insensitive. Matches a variant's own accession too. |
plasticdb_id | Exact PlasticDB v1 identifier. A variant matches through its parent. |
plastic · include_subtypes | Exact plastic name, abbreviation or full name, case-insensitive — PE or Polyethylene, never a substring, so PE does not match PET or LDPE. Add include_subtypes=true for its curated sub-types. The same rule on /microorganisms, /papers and every export. |
microorganism · lineage | Substrings: of the source organism's name or any of its synonyms, and of its taxonomic lineage string (lineage=Pseudomonadota). On an engineered variant these resolve through its parent protein. |
year · min_year · max_year | The year the record was first reported (its earliest supporting paper): exactly year, or within the range. |
substrate_form · evidence_method | Vocabulary: the physical form of the substrate, and the assay the evidence rests on. Exact names from the vocabularies; an unknown value is a 422 listing the allowed ones. See below. |
min_kcat · max_km · min_kcat_km · min_vmax | Kinetic thresholds (kcat in s⁻¹, Km in mM, kcat/Km in M⁻¹·s⁻¹), matched against the record's own measurements: a variant on its own kinetics, a wild type on its own and never on its variants'. Stored units are normalised before comparing — a kcat reported in min⁻¹ is divided by 60 — and a measurement whose unit is not recognised is left out rather than compared as if it were s⁻¹. A threshold must be a finite number; NaN or inf is a 422. |
min_tm · max_tm | Melting temperature range (°C) from the stored enzyme properties. Both bounds apply to the same measurement; a variant is judged on its own Tm. |
has_structure | true: the record has its own experimentally determined structure with a PDB id (predicted AlphaFold/ESMFold models do not count). false: it has none. |
has_sequence | true / false: the record does / does not carry an amino-acid sequence. |
ec_number | EC number prefix, matched on whole levels: 3.1.1 matches 3.1.1.101 and 3.1.1.-, never 3.1.10.x. A variant inherits its parent's EC number. |
polarity · include_negative | Vocabulary (positive, negative, inconclusive): records with at least one result of that polarity. Positive-only by default; include_negative=true drops the result filter and returns every public record. |
kind | wild_type, engineered or both (default). See below. |
sort_by · sort_order · limit · offset · page | Ordering and paging. limit defaults to 100 and caps at 500. sort_by is one of year, name, accession, id (plus plasticdb_id, genbank_id with kind=wild_type); anything else is a 422. See Paging. |
Engineered variants
PlasticDB treats an engineered variant as a first-class record with its own accession, its own kinetics and its own papers. To walk one enzyme's engineering lineage:
curl https://v2.plasticdb.org/api/v1/engineering/proteins/PBG5RN
{ "accession": "PBG5RN", "name": "IsPETase",
"variants": [ { "id": 8, "accession": "PBM3M6", "mutations": [ { "raw": "N233K", ... } ], ... } ],
"wild_type_kinetics": [ ... ], "wild_type_properties": [ ... ] }
One variant in full — mutations, engineering method, kinetics, structures, parent — by
its accession or its numeric id (/variants/8 is the same
record). Its own sequence and structure are at /variants/{accession}/fasta
and /variants/{accession}/structure.
curl https://v2.plasticdb.org/api/v1/variants/PBM3M6
Or list every enzyme that has been engineered at all:
curl "https://v2.plasticdb.org/api/v1/engineering/proteins?limit=100"
Wild-type and engineered are one list, selected with kind
/proteins and the protein exports describe a union of two entity types:
kind | Returns |
|---|---|
both (default) | Wild-type proteins followed by engineered variants. |
wild_type | Wild-type proteins only. |
engineered | Engineered variants only. |
Any other value is a 422, so a typo tells you rather than quietly returning the wrong dataset.
Every row carries a kind column. Because the two entity types have independent
id sequences, the row key is (kind, id) — id
alone is not unique in a kind=both result. Variant rows also carry
parent_accession.
Every filter applies to both entity types. A variant resolves one of two
ways. Inherited: microorganism, lineage and
plasticdb_id have no column on a variant, so they resolve through its parent —
?microorganism=Pseudideonella&kind=engineered returns the
Pseudideonella-derived variants. Its own: substrate_form,
evidence_method, polarity and the kinetic thresholds match the
variant's own experiments, never its parent's — ?min_kcat=5&kind=engineered
selects variants whose own kcat clears 5 s⁻¹, which is the point of having
engineered them.
Microorganisms
Organisms are addressed by their NCBI taxonomy id or by their PlasticDB accession (any
case) — /microorganisms/1547922 and /microorganisms/PB3E8R are
the same record:
curl https://v2.plasticdb.org/api/v1/microorganisms/1547922
{ "id": 2, "name": "Pseudideonella sakaiensis", "tax_id": 1547922, "accession": "PB3E8R",
"degradation_links": [ { "plastic": { "name": "PET" }, "result": "positive",
"methods": [ "hplc", "sem", "enzyme_assay", ... ],
"positive_count": 8, "negative_count": 0,
"positive_papers": 3, "negative_papers": 0, "paper_count": 3 } ],
"lineage": "...", "environments": [ ... ], "locations": [ ... ], "papers": [ ... ] }
*_count tallies interactions (one per strain or isolate a paper tested);
*_papers counts the distinct papers behind them, which is the number the
site's badges show. Expression hosts, knockouts and engineered strains
(strain_occurrence.genetic_background) are never credited to the species;
an organism record lists them under construct_interactions.
Everything reported to degrade PET:
curl "https://v2.plasticdb.org/api/v1/microorganisms?plastic=PET&limit=500"
| Method | Endpoint | What it gives you |
|---|---|---|
GET | /microorganisms | Filter by q, name, plastic, plastic_id, tax_id, temp_class, environment, location, year/min_year/max_year, lineage, the taxonomic ranks domain…genus, evidence_method, substrate_form, sole_carbon_source, aerobicity, polarity, include_negative, evidence_scope. Sort by name, tax_id or growth_temp_opt or id.
Exact: plastic (see above; include_subtypes=true widens it to sub-types), plastic_id, tax_id, year (first reported).
Substring (case-insensitive): name (the name or any synonym), environment (the isolation-environment code, e.g. marine matches marine_sediment and marine_water), location (country or region), lineage and the rank filters (genus=Pseudomonas matches genus:Pseudomonas in the lineage).
Vocabulary (422 for anything else): temp_class, evidence_method, substrate_form, sole_carbon_source, aerobicity, polarity, evidence_scope. |
GET | /microorganisms/{tax_id or accession} | One organism: lineage, plastics with per-method evidence counts, proteins, papers. Takes evidence_scope (organism_or_enzyme or whole_organism). |
GET | /microorganisms/accession/{accession} | The same record by PlasticDB accession, kept for existing links. An organism with no NCBI tax id is reachable by accession on either route. |
GET | /microorganisms/taxonomy/{rank}/{name} | Everything at one point in the taxonomy, e.g. /taxonomy/genus/Bacillus. A genus or species that has been renamed and has no organism filed under it any more answers 302 to its current name (genus/Ideonella → genus/Pseudideonella); each top_microorganisms entry carries evidence_count, which is 0 for a member with no degradation result recorded. |
GET | /microorganisms/search?q= | Type-ahead by name or tax id; up to 25 matches. |
GET | /microorganisms/names | Every name and id, for populating a local lookup. |
GET | /microorganisms/lineage-levels | Distinct values per taxonomic rank, for building filter menus. Narrow with domain, kingdom, phylum, class, order. |
evidence_scope decides whether an organism is credited with a plastic through
its own assays only, or also through its enzymes. It defaults to the inclusive
organism_or_enzyme; pass whole_organism for the strict reading, or
enzyme_only for organisms credited solely through an enzyme. There is no
value called organism. enzyme_only selects organisms, so
it is accepted by /microorganisms and its exports only; the organism detail,
/plastics, /plastics/{id}, /stats and
/stats/charts take the first two and answer 422 for anything else.
Assay conditions. evidence_method (a method-vocabulary name such as
co2_evolution), substrate_form (e.g. film, powder),
sole_carbon_source (yes, no, not_stated) and
aerobicity (aerobic, anaerobic, microaerophilic,
not_reported — which also matches experiments with no value) must all hold in the
same experiment, on an interaction that also satisfies plastic and
polarity and is credited under evidence_scope. So
?plastic=LDPE&evidence_method=co2_evolution&sole_carbon_source=yes means a
positive CO₂-evolution assay on LDPE in which LDPE was the sole carbon source.
temp_class (psychrophile <20 °C,
mesophile 20–45 °C, thermophile ≥45 °C) buckets the organism's
recorded optimum growth temperature (growth_temp_opt), not any assay
temperature; organisms with no recorded optimum match none of the three. An unknown value
(e.g. thermophilic) is a 422, as are unknown sole_carbon_source and
aerobicity values.
Plastics, papers, compounds and pathways
| Method | Endpoint | What it gives you |
|---|---|---|
GET | /plastics | Every plastic in the vocabulary with its counts of degrading organisms and enzymes (distinct actors with a positive result). Filter with q (full-text), name, cas_number and polymer_class (substrings, case-insensitive — name=PET also returns PETG and PET blends; use plastic_id for one record), parent_id, and has_evidence (true: at least one degrader; false: none recorded). Sort by name, full_name, cas_number, microorganism_count or protein_count. |
GET | /plastics/{plastic_id} | One polymer with its degraders and enzymes, and paper_count — the total /papers?plastic=<name> returns. Page the sub-lists with microorganism_offset/microorganism_limit and protein_offset/protein_limit (25 each by default, at most 500). microorganism_polarity/protein_polarity (positive, inconclusive, negative) pick which bucket is paged; microorganism_sort_by (papers, result, name) and protein_sort_by (also accession, plasticdb_id, id) order it. Unknown values are a 422. |
GET | /plastics/names | Names and ids only, for dropdowns. |
GET | /papers | Literature. Filter with q (full-text), year (exact) or min_year/max_year, doi, journal, authors and country (substrings, case-insensitive), and by the evidence a paper reports:
plastic (exact name, abbreviation or full name — PE never matches PET; add include_subtypes=true for curated sub-types),
microorganism (NCBI tax ID, accession or name; includes enzymes purified from it),
protein (accession), evidence_method (a method name),
result (positive, negative, inconclusive) and
has_quantitative=true (at least one numeric observation).
The evidence filters all apply to one reported result, so ?plastic=PE&evidence_method=weight_loss means a weight-loss assay on PE.
An unknown evidence_method or result is a 422.
limit defaults to 20 here, not 100. |
GET | /papers/{paper_id} · /papers/{paper_id}/footprint | One paper, and an aggregate of every record it contributed. |
GET | /compounds · /compounds/{compound_id} | Monomers, additives and pathway intermediates. Filter with q and is_additive. |
GET | /pathways · /pathways/{pathway_id} | Curated degradation pathways and their reaction steps. |
GET | /pathways/by-plastic/{plastic_id} · /by-plastic/{plastic_id}/graph | The pathways for one polymer, and the same thing shaped as a node/edge graph. [] means the plastic has no curated pathway; a plastic id that does not exist is a 404. |
GET | /stats · /stats/charts · /stats/by-country | Database totals, chart series and the country map. |
GET | /vocabularies | Every controlled vocabulary in one response — methods, measurement types, environments, substrate forms, variant types, plastics, compounds. Cached an hour. See Controlled Vocabularies. |
# PET, with its first 25 degraders and 25 enzymes curl https://v2.plasticdb.org/api/v1/plastics/16 # the degradation pathways curated for PET curl https://v2.plasticdb.org/api/v1/pathways/by-plastic/16
Read allowed values from /vocabularies at runtime rather than hard-coding
them, so a client keeps working as the vocabularies grow.
Search
curl "https://v2.plasticdb.org/api/v1/search?q=PET+hydrolase&limit=20"
{ "total": 37, "items": [ { "id": 292, "name": "Hong, H., Ki, D., ...", "entity_type": "paper",
"rank": 1.0, "extra": "Year: 2023" } ] }
| Method | Endpoint | What it gives you |
|---|---|---|
GET | /search?q= | One ranked list across microorganisms, proteins, plastics and papers. Each hit carries entity_type and route_key, which together address the detail record. limit defaults to 20, caps at 500. |
GET | /search/grouped?q= | The same matches bucketed by entity type instead of interleaved. limit is per bucket, defaults to 10 and caps at 100. |
GET | /search/autocomplete?q= | Top five per entity type for type-ahead. q must be at least 2 characters. Allowed 300 requests/hour rather than 200. |
Search is full-text plus fuzzy trigram matching, so a misspelling still finds the record.
To export a result set rather than page through it, see
/downloads/search.zip below.
Bulk downloads
The export endpoints under /api/v1/downloads are public and return the whole
database or any filtered slice of it. They accept exactly the same filter
parameters as the list endpoint they mirror, so an export always matches the
result set you were looking at.
# every wild-type PET enzyme, as FASTA, ready for a DIAMOND database curl -OJ "https://v2.plasticdb.org/api/v1/downloads/proteins.fasta?plastic=PET&kind=wild_type" # the whole database in one archive curl -OJ https://v2.plasticdb.org/api/v1/downloads/dataset.zip
| Endpoint | Contents |
|---|---|
/downloads/dataset.zip · /dataset.json | The complete dataset, including protein_variants and the evidence spine — interactions, experiments and observations, keyed to each other and to the entity tables. The zip carries README.md (every table and how they join) and LICENSE.txt (CC BY 4.0); dataset.json opens with a metadata object carrying the same licence and citation. Takes no filters. |
/downloads/proteins.fasta · /proteins.csv · /proteins.json | Enzymes — wild-type and engineered, selected with kind. Accepts the whole /proteins filter set. |
/downloads/microorganisms.csv · .json | Organisms. Accepts the whole /microorganisms filter set. |
/downloads/plastics.csv · .json | Polymers. Filters: q, name, cas_number, plastic_id, polymer_class, parent_id. |
/downloads/papers.csv · .json | Literature (also served as /references.*). Accepts the /papers filters, including the evidence filters. |
/downloads/protein-variants.csv · .json | Engineered-variant detail: mutations, engineering method, kinetics, sequence (also served as /engineered-versions.*). Filters: parent_protein_id, parent_plasticdb_id, variant_type, engineering_method. |
/downloads/search.zip · /search.json | A search result set, exported. Requires q. |
Joining a record to its sequence
Every row carries seq_id, which is exactly the identifier of that record's
FASTA entry, and it is never empty. Join on seq_id —
proteins.csv ↔ proteins.fasta ↔
protein-variants.csv all agree on it. accession is the real
registry value and the better key when you specifically want registered records, but it is
empty for a record whose accession has not been minted yet; seq_id falls back
to PlasticDB_{id} (or PlasticDB_var_{id} for a variant) so those
records stay joinable.
Variant rows populate their own interactions, kinetics, structures and papers; the source
organism is inherited from the parent. Columns with no variant analogue
(plasticdb_id, gene, genbank_id,
uniprot_*, …) are empty.
Two deliberate differences. (1) proteins.fasta is a sequence file, so it omits records with no sequence; proteins.csv/.json keep them as metadata rows and can be one or more rows longer. (2) dataset.zip/dataset.json are a complete archive and apply no include_negative filter, so their proteins set is a superset of the endpoint's default — pass ?include_negative=true to match.
Selecting evidence by assay method
/proteins, /microorganisms and their exports accept
evidence_method, matching the assay a record's evidence rests on —
weight_loss, hplc, sem, clear_zone,
ftir and so on, matched exactly against the method vocabulary.
curl -OJ "https://v2.plasticdb.org/api/v1/downloads/proteins.fasta?evidence_method=hplc"
This is a descriptive axis, not a quality ranking. Every record in PlasticDB is curated from published literature; no method ranks above another, and a record is not more or less trustworthy for having been assayed one way rather than another. Pick the assay that suits your analysis.
Checking a download before you pull it
HEAD works on every URL on this site, exports included: same status, same
Content-Type and Content-Disposition, no body — so
curl -I, link checkers and uptime monitors all work. No
Content-Length is advertised, because exports are streamed and their size is
not known until they are generated; a 0 there would be a lie rather than a
number you could use. HEAD is deliberately absent from the OpenAPI schema:
RFC 9110
defines it as GET minus the body, so listing a twin of every endpoint would
double the reference without telling you anything new.
Which version of the data did you use?
PlasticDB publishes weekly. Curation runs continuously, but approved
records stay invisible until a release makes them public — every Thursday at 06:20 UTC.
The dataset version is the ISO year and week, such as 2026.39 (an extra
out-of-cycle release in the same week gets a suffix, 2026.39.2), and it is
deliberately not the application version in the page footer:
1.32.0 says which code is running, the dataset version says which
data you searched.
curl https://v2.plasticdb.org/api/v1/releases/current
{
"version": "2026.39",
"released_at": "2026-09-24T06:23:11Z",
"doi": null,
"snapshot_release_label": null,
"citation": "PlasticDB dataset release 2026.39. …",
"license": "CC BY 4.0", ...
}
The shape, not live values: ask the endpoint for today's. doi and
snapshot_release_label are null until a release has been
archived with a DOI.
Record version in your methods section, and quote
citation verbatim. That is the whole point of the field: it turns
"downloaded from PlasticDB in August 2026" into a statement someone else can reproduce.
| Method | Endpoint | What it gives you |
|---|---|---|
GET | /releases/current | The live release: version, date, DOI, ready-made citation. Every field is null before the first release. |
GET | /releases | Version history, newest first. limit defaults to 50 and caps at 200. |
GET | /releases/{version} | One release and what it published — counts of promoted enzymes and organisms, timestamps. |
GET | /snapshots · /snapshots/latest · /snapshots/{id} | The archived artefacts themselves: release_label, entity_counts, figshare_doi, checksum, file_size. /snapshots/latest is a 404 until a snapshot has been published. |
doi and snapshot_release_label may name an
earlier release than version — that is correct, and it means the archive you
cite genuinely contains the data you used. Full cadence and reasoning:
docs/dataset-releases.md.
The exports under /downloads always serve the current release; there
is no way to ask them for an older one. To pin an analysis to a past release, cite the
Figshare DOI of that release's snapshot.
Run an analysis
Everything above is anonymous. This section is not: a job is owned by whoever launched it, so from here on you need a credential — one header, described in Authentication below. Get an API key once, then read on.
The fastest thing: search one sequence, synchronously
POST /bioinformatics/search-sequence runs DIAMOND against the PlasticDB
protein database and returns the hits in the HTTP response. No job, no polling.
Its parameters go in the query string, not in a JSON body.
curl -X POST -H "X-API-Key: $PLASTICDB_KEY" \ "https://v2.plasticdb.org/api/v1/bioinformatics/search-sequence?sequence=MNFPRASRLMQAAVLGGLMAVSAAATAQ...&evalue=1e-5&pident=30"
{ "hits": [ { "subject_seq_id": "PBG5RN", "protein_name": "IsPETase", "entity_type": "protein",
"pident": 99.7, "length": 290, "evalue": 1.2e-180, "bitscore": 601.2,
"qstart": 1, "qend": 290, "sstart": 1, "send": 290 } ] }
A sequence long enough to overflow your proxy's URL limit will not fit in a query string. For anything larger than a single protein, upload it as a job instead.
Submitting a job
POST /jobs takes a multipart upload — the same path the web tools use:
curl -X POST https://v2.plasticdb.org/api/v1/jobs \ -H "X-API-Key: $PLASTICDB_KEY" \ -F "job_type=annotate_genome" \ -F "file=@my_genome.faa" \ -F "evalue=1e-5" -F "pident=30" -F "blast_type=blastp"
{ "id": "5f2c…", "job_type": "annotate_genome", "status": "waiting", "result_count": 0, ... }
job_type | What it does | Upload |
|---|---|---|
annotate_gene | Align one or a few sequences against PlasticDB. | file (FASTA) |
annotate_genome | Screen a whole proteome or genome for plastic-active enzymes. | file (FASTA) |
annotate_list | Annotate a TSV list of organism names. | file (TSV) plus genus_column, species_column, field_separator |
compare_genomes | Compare the plastic-degrading repertoire of several genomes. | files — two or more, repeat the field |
pathway_analysis | Map hits onto curated degradation pathways. | file (FASTA) |
hmm_screen | Search profile HMMs built from families of curated plastic-active enzymes — reaches remote homologs a pairwise alignment misses. | file (FASTA, protein) plus optional target_polymers |
One pipeline is not reached this way. The Taxonomy Tree is POST /tree/build, which
takes a microorganism filter or a list of ids rather than a sequence upload, and needs no
credential. It is deliberately not a job_type here: this endpoint
saves whatever you upload as the job's input FASTA, and the tree builder never reads it,
so a build_tree submission used to return 201 and then build
something unrelated to what was sent (#459).
target_polymers applies to hmm_screen only — repeat the field to
send several, and omit it to summarise every polymer. It narrows the per-polymer tally on
the job; every hit is stored either way. An unrecognised name is a 422 listing
the accepted set. A hit is a profile match, not a measurement: it says a sequence resembles
a family whose members have been reported acting on a polymer, not that it degrades one.
Parameters, all optional: evalue (default 1e-5, must be in the range 0 < e ≤ 10), pident (default 30, 0–100), blast_type (blastp or blastx), organism_type (no_signalp, arch, gram+, gram-, euk) and include_variants (default false — off so results are natural degraders rather than a protein's own mutants; it is recorded on the job so a result set can be read months later without guessing which database produced it).
Getting the results back
curl -H "X-API-Key: $PLASTICDB_KEY" https://v2.plasticdb.org/api/v1/jobs/5f2c… curl -OJ -H "X-API-Key: $PLASTICDB_KEY" https://v2.plasticdb.org/api/v1/jobs/5f2c…/download
| Method | Endpoint | What it gives you |
|---|---|---|
POST | /jobs | Submit a job. 30/hour. |
GET | /jobs | Your jobs, newest first. |
GET | /jobs/{job_id} | Status and results, including the DIAMOND hits. |
GET | /jobs/{job_id}/download | The results file — a TSV, or for compare_genomes a ZIP of comparison_matrix.tsv plus one hits TSV per genome. |
GET | /jobs/{job_id}/comparison?format=tsv|csv | compare_genomes only: the genome × plastic matrix as a file (404 for any other job). |
Each hit's protein carries the matched record's source_microorganism
(null when none is on record or it is not yet public), its own degradation_links
per plastic with result counts (methods is left empty here — read it from
/proteins/{accession}), and paper_count. The paper list itself is not
embedded per hit; /proteins/{accession} has it.
A finished compare_genomes job also returns comparison: the plastics
(columns) and one row per genome with counts (query sequences with a hit to a
record that has a positive result for that plastic, each sequence once per plastic),
normalised (per 1,000 input proteins for blastp, per megabase for
blastx; null for jobs run before genome sizes were recorded),
positive, and not_positive — sequences whose hits carry no positive
result for any plastic, which are never counted as potential. It is built from the hits and
their current evidence each time it is read.
status runs waiting → running → done,
or error with an error_message. Poll the detail endpoint. Two
fields tell you where the output is: result_count counts the structured hits
in the response, and has_result_file says whether
/download will return anything —
annotate_list and pathway_analysis write a file instead of
structured rows, so their result_count is 0 even when the job
found plenty.
A job you launched is yours: GET /jobs/{job_id} and its
/download answer 403 to anyone else. Jobs launched anonymously
from the web tools have no owner and are readable by anyone holding the id, which is why
the read routes accept a request with no credential at all.
The JSON job endpoints
/bioinformatics/* mirrors the same pipelines without a multipart upload,
taking base64 FASTA in the query string — built for scripted and agent access. All require
a credential and are limited to 30/hour.
| Endpoint | Required parameter |
|---|---|
POST /bioinformatics/annotate-gene | fasta_b64 — the FASTA file, base64-encoded. Plus the same evalue, pident, blast_type, organism_type, include_variants as above. |
POST /bioinformatics/annotate-genome | |
POST /bioinformatics/annotate-list | |
POST /bioinformatics/pathway-analysis | |
POST /bioinformatics/compare-genomes | fasta_b64 repeated, once per genome (at least two), plus optional labels repeated in the same order to name them (default genome_1, genome_2, …). Names must be distinct. |
POST /bioinformatics/hmm-screen | fasta_b64, plus target_polymers (repeat the parameter; omit to screen every profile). |
POST /bioinformatics/search-sequence | sequence — a plain sequence string. Synchronous; returns hits, not a job. |
Each returns the same job object as POST /jobs, so poll
/jobs/{job_id} the same way — except /search-sequence, which
answers with the hits themselves. The Taxonomy Tree builder is separate:
POST /tree/build takes a microorganism filter or an explicit id list, works
without a credential, and the result is read from GET /tree/{job_id}: an
NCBI-taxonomy cladogram (ranks only, no branch lengths) in result_newick,
with the per-organism plastic matrix and the counts of organisms left out in
result_metadata.
Authentication
You need this for exactly three things: launching jobs, submitting data, and your own account. Nothing above Run an analysis requires it.
Get an API key
Register, confirm your email, then generate a key at Profile › API Keys. It is shown once and never expires until you revoke it, which is what makes it the right credential for a script or a pipeline. Send it as a header:
curl -H "X-API-Key: $PLASTICDB_KEY" https://v2.plasticdb.org/api/v1/jobs
Or exchange your password for a short-lived JWT and send that as a bearer token:
curl -X POST https://v2.plasticdb.org/api/v1/auth/login \
-d "username=you@example.org" -d "password=…"
# → { "access_token": "eyJhbG…", "token_type": "bearer" }
curl -H "Authorization: Bearer eyJhbG…" https://v2.plasticdb.org/api/v1/auth/me
An API key is also accepted as a bearer token, so a single
Authorization: Bearer code path works for both if you prefer one.
Being logged in to the website does not authenticate an API call
The browser session is an httpOnly cookie, and the API
deliberately refuses it — a cookie is attached by the browser automatically,
so accepting one would make every API route reachable by a hostile page on another
site (CSRF). The pages you click through use the cookie; the API uses a header you have
to set on purpose. Practically: a fetch() from a page must send
Authorization: Bearer …, and a plain <a download> link
cannot authenticate at all, because a link cannot set a header.
| Method | Endpoint | Auth | Description |
|---|---|---|---|
POST | /auth/register | Public | Create an account, inactive until email confirmation. Body: {email, password, full_name}. 5/hour. |
POST | /auth/login | Public | Get a JWT. Form body: {username: email, password}. Returns {access_token, token_type}. 10/minute. |
GET | /auth/confirm-email?token= | Public | Confirm a registration email. |
POST | /auth/resend-confirmation | Public | Resend the confirmation email. 3/hour. |
POST | /auth/forgot-password · /auth/reset-password | Public | Password-reset request and completion. 5/hour each. |
POST | /auth/logout | Public | Clear the browser session cookie. Irrelevant to a header-authenticated client. |
GET · PATCH | /auth/me | Auth | Read or update your profile. PATCH body: {full_name, password}. |
POST · DELETE | /auth/api-key | Auth | Generate a new API key — returned once, save it — or revoke the current one. |
Contribute data
Submissions are curated before they appear, and are published by the next weekly release. The record shape is documented at Contributing Data.
| Method | Endpoint | Auth | Description |
|---|---|---|---|
POST | /submissions | Auth | Submit interactions with their actors, experiments and observations. Requires "license_accepted": true: accepted records are published under CC BY 4.0. 30/hour. |
GET | /submissions/mine | Auth | Your submissions and their curation status. |
GET | /submissions/{id}/matches | Auth | Existing records your submission may duplicate. |
POST | /submissions/lookup-doi | Public | Paper metadata by DOI — the database first, then Crossref. Body: {doi}. 30/hour. |
GET | /submissions/lookup-taxid?name= | Public | Search NCBI taxonomy by name. 30/hour. |
POST | /submissions/lookup-sequence | Auth | Fetch a sequence by accession. Body: {accession_id, database: genbank|uniprot}. 30/minute. |
POST | /submissions/validate-sequence | Auth | Check whether a string is a valid amino-acid or nucleotide sequence. 30/minute. |
GET | /submissions/isolation-environments · /sample-types · /evidence-types · /methods | Public | Vocabulary helpers for building a submission. |
Reference
Paging and identifiers
List endpoints return { "total": …, "items": [ … ] } and page with
limit and offset. limit caps at 500
almost everywhere — the exceptions are /releases and /snapshots
(200) and /search/grouped (100). Defaults differ: 100 for
/proteins, /microorganisms, /plastics,
/compounds and /pathways; 50 for /jobs,
/releases and /snapshots; 20 for /papers and
/search; 10 for /search/grouped.
Paging is stable. Every sort ends with a unique key (the record id), so
walking offset=0, limit, 2×limit, … returns each record exactly once however
many records share a year or a name. page (1-based) is accepted as shorthand
for offset=(page−1)×limit; send page or a non-zero
offset, not both (422). sort_by and sort_order take
a fixed set of values, listed per endpoint in the OpenAPI schema; an unknown one is a
422 naming the allowed values, never a silent fallback to the default order.
Sorting. Text sorts ignore case (uncultured bacterium sorts
among the Us, not after Z), and records with no value for the
sort key come last in both directions — so
sort_by=growth_temp_opt&sort_order=desc opens with the hottest organism, not
with the hundreds that have no recorded temperature.
Identifiers. Dates are ISO 8601. Proteins, engineered variants and
organisms all carry a PlasticDB accession (PB + 4 characters),
and it works on every route for its record, in any case: /proteins/{accession}
(and /fasta, /structure), /variants/{accession},
/microorganisms/{accession}. Those routes also still take the numeric id (an
organism's NCBI tax_id). Plastics, papers, compounds and pathways have
integer ids only; jobs and users are UUIDs.
Unknown query parameters are ignored, not rejected. A misspelled
parameter name (?plastc=PET) is dropped and the request answers as
if it had not been sent — a misspelled value of a known parameter is a
422. This is deliberate: rejecting unknown names would break every client
that sends a parameter a later release renames or retires. If a filter seems to have no
effect, check its spelling against the OpenAPI schema.
Rate limits
Limits are set per route, and only on the routes listed below; every other
route — the detail routes for plastics, microorganisms, proteins, papers and compounds,
/stats, /vocabularies, /releases,
/snapshots — is not rate-limited at all. Each listed route keeps its own
count, so 200 requests to /proteins do not use up /papers.
Who shares a count. A request carrying a login token
(Authorization: Bearer with the JWT from /auth/login) is counted
for that user. Every other request — anonymous, or authenticated with an
API key (in X-API-Key or as a bearer token) — is counted for
its IP address, so everyone behind one institutional gateway shares it.
Over the limit is a 429 with
{"detail": "Rate limit exceeded: …"}; wait for the window to roll over and
retry. No rate-limit headers are sent, so pace a script yourself — or use the
bulk downloads, which return a whole filtered set in one request.
| Endpoint(s) | Limit |
|---|---|
/proteins, /microorganisms, /plastics, /papers, /compounds (the lists, not the detail routes), /search, /search/grouped, every /downloads/* export, /variants/*, /engineering/*, /pathways/by-plastic/*/graph | 200 / hour |
/search/autocomplete | 300 / hour |
POST /jobs, all POST /bioinformatics/*, POST /tree/build, POST /submissions, POST /annotations, POST /submissions/lookup-doi, GET /submissions/lookup-taxid | 30 / hour |
POST /submissions/lookup-sequence, /validate-sequence | 30 / minute |
POST /auth/login | 10 / minute |
POST /auth/register, /forgot-password, /reset-password | 5 / hour |
POST /auth/resend-confirmation | 3 / hour |
Status codes
| Code | Means |
|---|---|
301 | This accession was merged into another record, or it belongs to an engineered variant and the record is at /variants/{accession}. Follow the Location header. |
302 | A taxonomy name that is now a synonym (see /microorganisms/taxonomy/{rank}/{name}). Location is the current name. Temporary on purpose: the mapping follows the data. |
401 | No credential, or one the server does not recognise. See Authentication — a browser session cookie is not a credential here. |
403 | Authenticated, but the resource is someone else's — most often another user's job. |
404 | No such record, or the record is not public. |
410 | This accession was withdrawn. It will not come back, and it has no successor. |
422 | A parameter value is out of range or misspelled — ?kind=enginered, ?evidence_scope=organism, evalue=0. Deliberately an error rather than a silent fallback to the default, which would hand you the wrong dataset with a 200. |
429 | Rate limited. |
Error bodies
Every error is a JSON object with a detail key. On a 422 from a
read (GET) endpoint, detail is always a list, one entry per
problem, each naming the parameter in loc:
{ "detail": [ { "type": "enum", "loc": ["query", "evidence_scope"],
"msg": "Input should be 'organism_or_enzyme' or 'whole_organism'",
"input": "organism" } ] }
Every other status carries a human-readable string: {"detail": "Protein not
found"}. The POST job and submission routes answer some of their own
checks (an empty upload, a bad e-value) with a string detail too; read
detail as "a list of problems, or one sentence" and you will handle both.
Not covered here
The curation, ingestion, annotation, forum, chat and administrative routes are site machinery rather than a research interface, and most require a curator or superuser account. They are all in the generated schema if you need them.