Bioinformatics Tools

Analyse your own sequences and taxa against PlasticDB.

Overview

PlasticDB provides several analysis pipelines. Most align your input against the PlasticDB enzyme database with DIAMOND (a fast BLAST alternative) and, where relevant, predict signal peptides with DeepSig, which runs on PlasticDB's own servers. One tool instead matches taxonomy names, one searches profile hidden Markov models with HMMER, and one draws a taxonomy tree. Jobs run asynchronously — you can close the page and check back later under My Jobs.

Running a job requires a (free) account, except Taxonomy Tree, which is open to everyone. All tools are reachable from the Tools menu.

Sequence-based tools (DIAMOND)

Annotate Gene

/tools/annotate-gene — upload a FASTA file (there is no paste box) holding one or a few protein or nucleotide sequences to find their closest PlasticDB matches. Supports BLASTP (protein) and BLASTX (nucleotide, translated in all six frames). For BLASTP searches you can also ask for signal-peptide prediction (see Signal peptides below).

Annotate Genome

/tools/annotate-genome — upload a multi-FASTA file of predicted proteins from a genome. Each sequence is aligned independently against PlasticDB to surface every candidate plastic-degrading gene. Best for whole-genome annotation. The result is the table of hits (also as a TSV download); there is no per-genome summary.

Nucleotide input works with BLASTX, but no tool calls genes: BLASTX translates each contig as a whole, so you get one row per contig and match, with the coordinates only in the TSV. For assemblies, call genes first (e.g. Prodigal) and upload the predicted proteins with BLASTP.

Compare Genomes

/tools/compare-genomes — upload several FASTA files (one per genome or metagenome) and compare their enzyme-hit profiles as a matrix, broken down by plastic type. Useful for contrasting the plastic-degrading potential of different strains or communities.

Each cell of the matrix counts the query sequences in a genome with at least one hit to a PlasticDB record that has a positive result of its own for that plastic; a sequence is counted once per plastic however many records it hits. Hits to records whose only results are negative or inconclusive never add to a cell — each genome's count of such sequences is shown in a separate column. Counts are normalised per 1,000 input proteins (FASTA records) for BLASTP, and per megabase of input sequence for BLASTX, where a record is a contig or read rather than a gene. The job page draws the matrix as a heatmap and a per-genome bar chart, each with a table of the same numbers; the matrix downloads as TSV or CSV, and the zip download carries it alongside each genome's hits.

Pathway Analysis

/tools/pathway-analysis — upload a multi-FASTA file (proteins with BLASTP, or nucleotides with BLASTX) to map your sequences onto the Pseudideonella sakaiensis PET-degradation pathway; hits light up the steps they cover on a pathway diagram. Each step is judged against the curated enzyme plus every PlasticDB protein with positive evidence and the same EC number. A step is called present only when a hit has ≥40% identity, e-value ≤1e-10, and covers ≥70% of the reference enzyme and ≥50% of your protein (all four adjustable on the form); the results list every step with its EC number, best hit, identity and coverage, and name the missing steps. This tool is currently PET-specific, and the pathway ends at protocatechuate ring cleavage: ethylene glycol catabolism and the route to the TCA cycle are not curated yet.

Profile-based tool (HMMER)

HMM Screen

/tools/hmm-screen — upload a multi-FASTA file of protein sequences and search it against PlasticDB's profile hidden Markov models with hmmsearch. The profiles are built from aligned families of curated plastic-degrading enzymes rather than from single sequences, so they encode what a family has in common and reach further than a pairwise alignment does.

Open to every signed-in user, like the other upload tools. The profiles are built from PlasticDB's curated records, which are public; a handful may still be staged for the next weekly dataset release when a library is built. A profile is a statistical summary of a family alignment: it names no record and reproduces no sequence.

Use it rather than the DIAMOND tools when your sequences have no close homolog in the database — screening predicted proteins from a metagenome is the case it exists for. Optionally restrict the per-polymer tally on the results page to particular polymers; every hit is stored either way.

A hit is a similarity match to a profile, not a measurement. It says a sequence resembles a family whose members have been reported acting on a polymer; it does not say your sequence degrades anything. Treat every hit as a candidate to test. hmmsearch does not translate nucleotides, so call genes before uploading contigs. The equivalent API endpoint is POST /api/v1/bioinformatics/hmm-screen (see the API documentation), or POST /jobs with job_type=hmm_screen for a multipart upload.

Taxonomy-based tool

Annotate Taxa Table

/tools/annotate-list — upload a taxonomy table (for example DADA2/QIIME2 output) with genus and species columns. Rather than aligning sequences, this tool matches the taxon names against PlasticDB's own organism records: their current names, and the NCBI synonyms and lineages already stored on each record. It makes no live call to NCBI. It reports which of your taxa have records in PlasticDB, how each name matched (exact, synonym, fuzzy spelling within the genus, or genus only) and what evidence those records hold. Use it to scan microbiome or environmental survey data for plastic-degrading taxa. A name PlasticDB holds no record for comes back as No match, even if it is a valid NCBI taxon.

Taxonomy Tree

/tools/build-tree (public) draws a taxonomy tree of a filtered set of PlasticDB microorganisms, or of all of them, and opens it in an interactive viewer. The View as Tree buttons on the microorganism, protein and plastic lists build the same tree from the list's filters.

It is an NCBI-taxonomy cladogram: organisms are grouped by the ranks of their stored NCBI lineage (domain, phylum, class, order, family, genus). It is not a phylogenetic tree — no sequences are aligned or compared, and the branches have no lengths, so the distance between two organisms on the drawing says nothing about how related they are beyond the ranks they share. Each tip carries a coloured dot for each plastic PlasticDB holds a positive result for.

  • Left out: entries that are not organisms — NCBI's synthetic construct (taxid 32630), and placeholder names such as uncultured bacterium, unidentified …, unclassified …, any … metagenome, or a bare bacterium or fungus (the same rule Annotate Taxa Table uses to refuse a name). Organisms with no stored lineage are also skipped. The viewer's footer counts each group.
  • Viewer: pan and zoom, rectangular or radial layout, per-plastic toggles, label and dot size, and click a grey node to collapse or expand its clade.
  • Downloads: Newick (.nwk, the whole tree as built, topology only), SVG (the tree as drawn, with the plastic key, as a standalone vector file) and PNG (the same at 2× resolution).

Common parameters

  • E-value — significance threshold; lower is more stringent. Default 1e-5, on every tool page and in the API. The pages offer 1e-3 (permissive, for exploratory or distant-homolog searches), 1e-5, 1e-10, 1e-20 and 1e-50; the API accepts any value in (0, 10].
  • Percent identity — minimum amino-acid identity (DIAMOND --id). Default 30%. Raise to 60–70% for near-exact matches.
  • Organism type — which signal-peptide model to use: Gram-positive, Gram-negative, Archaea, Eukaryotic, or no prediction (see below).
  • BLAST type — BLASTP for protein input, BLASTX for nucleotide input (translated in six frames). The input must match: a file of nucleotide sequences in BLASTP mode (or sent to HMM Screen), or of protein sequences in BLASTX mode, is refused at submission with a message naming the records, rather than run to an empty result. A sequence counts as nucleotide when at least 90% of its residues are A, C, G, T, U or N.
  • Input — every tool takes an uploaded file; none has a paste box.
  • Compressed input — every upload control accepts gzip. Upload genome.fasta.gz as-is; it is stored compressed and read directly by the search engines, and results are labelled genome.

Search settings (fixed)

These are not exposed as options; they are what every DIAMOND job runs with.

  • DIAMOND 2.0.6, blastp or blastx, at DIAMOND's default sensitivity (no --sensitive flags).
  • Up to 25 database matches per query — DIAMOND's default --max-target-seqs 25. A query that matches more than 25 PlasticDB records reports the 25 best.
  • No query or subject coverage cutoff; filtering is by E-value and percent identity only.
  • Output is tabular (outfmt 6) with qlen and slen appended: qseqid, sseqid, pident, length, mismatch, gapopen, qstart, qend, sstart, send, evalue, bitscore, qlen, slen. The TSV has no header row.
  • HMM Screen runs HMMER 3.4 hmmsearch with the chosen full-sequence E-value (default 1e-5).

Signal peptides

In Annotate Gene and Annotate Genome, choosing an Organism type other than No signal peptide runs DeepSig (Savojardo et al., Bioinformatics 2018; GPL-3.0) on the query sequences that had a database hit — up to 50, strongest hit first. It runs on PlasticDB's own servers, so your sequences are not sent anywhere else. The results page shows, for each query, whether a signal peptide was predicted and the last residue of the peptide (the cleavage site).

  • Gram-negative, Gram-positive and Eukaryotic select DeepSig's three models.
  • Archaea: DeepSig has no archaeal model. Archaeal jobs use the Gram-positive model, the closest match, and the results say so.
  • Prediction needs protein sequences, so it runs for BLASTP searches only.

If the local predictor is ever unavailable, the service may fall back to EMBL-EBI's hosted Phobius service instead. Phobius has a single model and ignores the organism type, and the results page names whichever tool actually ran. The privacy notice explains what that means for your sequences.

Running jobs well

  1. Choose the right tool — Annotate Gene for a single sequence, Annotate Genome for whole genomes, Annotate Taxa Table for taxonomic surveys.
  2. Check your input format — multi-FASTA files need valid > headers; taxa tables need genus and species columns.
  3. Be patient — large files can take several minutes. Jobs run in the background; track them from My Jobs.
  4. Read results critically — a hit with low identity (30–40%) may not indicate genuine plastic-degrading activity. Cross-reference with the literature and the evidence on each PlasticDB match.

These tools are also available programmatically via the Jobs API — see the REST API page.