Article¶
Use article commands for literature retrieval by disease, gene, drug, and identifier.
Typical article workflow¶
- search a topic,
- choose an identifier,
- retrieve default summary,
- request indexing, full text, or annotations only when needed.
Search articles¶
By gene and disease:
By keyword:
By author:
Author filtering is an authorship constraint. The default --source all route
limits candidate search to backends with native author fields: Europe PMC and,
when the other selected filters are compatible, PubMed. Filters such as
--open-access or --no-preprints can narrow the plan further to Europe PMC.
You can select --source europepmc or --source pubmed directly; PubTator3,
Semantic Scholar, and LitSense2 reject --author instead of treating the name
as free text.
Tune keyword-bearing relevance:
biomcp search article -k "Hirschsprung disease ganglion cells" --ranking-mode hybrid --weight-semantic 0.5 --weight-lexical 0.2 --limit 5
By date:
By year range:
Exclude preprints when supported by source metadata:
Query formulation¶
Turn a natural-language literature question into two parts:
- Put a known gene, disease, or drug in
-g/--gene,-d/--disease, or--drug. Article--geneaccepts one nonempty symbol without whitespace; put extra concepts in--keyword. - Put mechanisms, phenotypes, outcomes, datasets, and other provider-neutral free-text concepts in
-k/--keyword. - Put a known author in
-a/--authorand a known journal in--journal; do not put PubMed or Europe PMC field grammar in-k/--keyword. - Keyword also rejects
gene:,disease:, anddrug:labels at the start or after whitespace/(. To keep prose literal, put a literal double-quote byte immediately before every reserved label, such asreview of "drug: safety"; typed MCP encodes it as"keyword":["review of \"drug: safety\""]. A whole-value quote does not protect a later label, and shell or JSON quoting alone does not add runtime quote bytes. - If the question is asking which gene, disease, or drug fits the evidence and you do not know the entity yet, do not guess a typed flag. Start with keyword-only article search or run
biomcp discover "<question>"first. - Question-format article terms are acceptable: PubMed ESearch cleans bounded filler words from unfielded gene, disease, drug, and keyword terms provider-locally, while query echoes and non-PubMed sources keep the original wording.
- Use
--type reviewfor synthesis questions, list-style questions, and dataset surveys.
Keyword-only searches can also return exact entity suggestions. When the whole
keyword exactly matches a gene, drug, or disease vocabulary label or alias,
BioMCP can add a typed get gene, get drug, or get disease follow-up in
See also, _meta.next_commands, and JSON _meta.suggestions[]. The
structured suggestion object includes command, reason, and sections.
Multi-concept phrases such as BRAF V600E or lung cancer immunotherapy do
not get direct entity suggestions. A whole exact gene-plus-protein keyword such
as MSH2 p.L341P is handled separately: article search remains unchanged, but
successful Markdown and JSON output add biomcp variant articles "MSH2
p.L341P" to the follow-ups. Typed gene, disease, or drug context does not
suppress this exact-variant pivot.
For the complete strict-identity, union-shortlist, batch, selected full-text/assets,
and conditional citation/reference sequence, run biomcp skill
exact-variant-literature.
biomcp --json variant articles "BRAF p.V600E" --verify-identity adds an
identity projection from captured provider evidence without changing normal
retrieval recall. --confirmed-only requires verification and filters the
verified candidate pool before ranking and pagination. Retrieval aliases remain
provenance, not observed article evidence. Confirmation requires a returned-PMID typed PubTator Gene/Variant linkage whose
exact HGVS and CorrespondingGene/NCBI Gene-ID facts agree, or an optional bounded
ClinGen LDH CAid/gene/PMCID annotation with one exact text/table selector. LDH only
observes already retrieved candidates; absent coverage is not negative evidence.
Relations and same-passage co-occurrence are not enough.
Identity observations include their provider, locator, linked gene, observed alias,
typed provider_linkage, and canonical captured-content hash;
verification artifact hashes in --debug-plan are versioned clinically relevant
response/content subset audit facts, not retrieval-cache keys. Verification
never fetches supplements automatically, and an incomplete verification is not a
confirmation: it keeps the result incomplete, truncated, and without a total. Frozen
fixtures provide release proof; live probes are diagnostics.
For agent loops, --session <token> lets JSON article search compare the
current keyword with the previous successful article keyword search for the
same local token. The token is not a secret; use a short non-identifying label
such as lit-review-1. When post-stopword term overlap is at least 60%,
BioMCP can add JSON-only _meta.suggestions[] fallbacks after exact entity
suggestions: prior batch article --mode compact, discover, and a date-narrowed retry when
available. Session baselines expire after 10 minutes. Markdown output is
unchanged. The local record stores the token, keyword and normalized terms,
bounded PMID set, and update time under the resolved BioMCP cache root. Opening
the store removes expired records; upstream providers apply their own separate
logging and retention policies.
BioMCP keeps the session directory and records private to the current OS user
and refuses linked files before reading or replacing them.
Use --no-cache when the invocation must not read or write either managed
request-state store; BioMCP rejects --no-cache --session before transport.
Known anchor only:
Known anchor plus mechanism or process:
Unknown-entity disease-identification question:
Known drug plus mechanism:
Dataset or method question:
Multi-source federation¶
Article search fans out to PubTator3, Europe PMC, and PubMed by default when
the filter set is compatible. Known gene, disease, drug, and keyword queries
participate in that route. Semantic Scholar can still join the same query when
the filter set is compatible. An author filter is intentionally narrower: the
default route limits candidate search to Europe PMC and compatible PubMed,
whose native author fields make every returned candidate an authorship match
rather than a lexical match.
Semantic Scholar and LitSense2 are also available as explicit single-source
routes with --source semanticscholar and
--source litsense2. Any explicit source selection stays on that source: it
does not silently enrich rows through Semantic Scholar or another metadata
provider. The response reports candidate and enrichment source plans. BioMCP merges duplicates across PMID,
PMCID, and DOI where possible. S2_API_KEY upgrades the Semantic
Scholar leg to authenticated requests at 1 req/sec; without it, BioMCP uses
the shared unauthenticated pool at 1 req/2sec. Search results are still
deduplicated by PMID when BioMCP can resolve one.
Default --sort relevance is mode-aware:
- Keyword-bearing queries default to
--ranking-mode hybrid, using0.4*semantic + 0.3*lexical + 0.2*citations + 0.1*positionwith the LitSense2-derived semantic signal. - Entity-only queries default to
--ranking-mode lexical, preserving the existing calibrated PubMed rescue plus lexical directness comparator. --ranking-mode semanticsorts the LitSense2-derived semantic signal first and falls back to the lexical comparator for deterministic ties.- Rows without LitSense2 provenance contribute
ranking.semantic_score = 0in semantic-aware ranking modes. --weight-semantic,--weight-lexical,--weight-citations, and--weight-positionretune the hybrid formula.
Markdown preserves the merged rank order. JSON search rows are compact by
default: available PMID/PMCID/DOI/arXiv/Semantic Scholar identifiers, title,
journal, date, citation count, primary source, and tri-state retraction state.
Use biomcp --json search article -g BRAF --limit 5 --full to restore detailed
rows with matched_sources, ranking, first_index_date, influential counts,
scores, and abstract snippets. Every search row, compact or --full, keeps a
240-byte abstract snippet; biomcp get article <pmid> returns the whole
abstract.
--sort date replaces relevance ranking rather than refining it. Compact JSON,
--full JSON, and Markdown all emit an in-band warning when date sort is used.
Use --source <all, pubtator, europepmc, pubmed, semanticscholar, litsense2>
to select one backend or keep the default federated search.
BioMCP caps each federated source's contribution after deduplication and before
ranking. Default: 40% of --limit on federated pools with at least three
surviving primary sources. Rows count against their primary source after
deduplication. Use --max-per-source <N> to override that cap, use
--max-per-source 0 for the default cap explicitly, and set it equal to
--limit to disable capping.
Default article search excludes confirmed retractions unless you pass
--include-retracted. Sources that do not expose retraction metadata still
participate in the search, and compact and --full JSON search rows keep the
tri-state contract: "is_retracted": true, false, or null.
--type, --open-access, and --no-preprints are backend-compatibility
constraints rather than universal filters across every article source.
--type on --source all uses Europe PMC + PubMed when --open-access and
--no-preprints are both absent. If you add --open-access or
--no-preprints, PubMed becomes ineligible and BioMCP surfaces the Europe
PMC-only note in markdown, JSON, and debug-plan output instead of silently
pretending the filter applies across every source.
To search a single backend:
biomcp search article -g BRAF --source pubtator --limit 5
biomcp search article -g BRAF --source europepmc --limit 5
biomcp search article -g BRAF --source pubmed --limit 5
To force a tighter federated balance:
Get an article¶
Supported IDs are PMID (digits only), PMCID (e.g., PMC9984800), and DOI
(e.g., 10.1056/NEJMoa1203421). Publisher PIIs (e.g., S1535610826000103) are not
indexed by PubMed or Europe PMC and cannot be resolved.
Detail returns every author name supplied by the selected metadata source in
source order. JSON includes authors, author_count (the number returned),
author_completeness (complete, source_limited, or unavailable), and
author_source (pubtator or europepmc). PubTator's structured author list is
complete; Europe PMC's display string is source_limited, so BioMCP does not
claim it reconstructs a ground-truth author list. Markdown prints the same names
and status without inserting an ellipsis or a synthetic author.
Default article output can include an optional Semantic Scholar section with
TLDR text, influence counts, and open-access PDF metadata when that paper
resolves in Semantic Scholar. S2_API_KEY makes those requests authenticated;
without it, BioMCP uses the shared pool. search article --source supports
all, pubtator, europepmc, pubmed, semanticscholar, and litsense2;
Semantic Scholar remains an automatic compatible leg and can also be queried
alone with --source semanticscholar.
Request specific sections¶
Full text section:
This uses the default article full-text ladder: XML first, then PMC HTML when the XML path misses for a PMCID-backed article. It never falls back to PDF. BioMCP accepts a winner only when JATS or PMC HTML structure contains a nonblank article-body block; title/front/abstract-only and metadata-only responses remain healthy partial results and the ladder continues. A source abstract fills a missing or blank article abstract but never replaces a nonblank base abstract.
When full text resolves, BioMCP prints a local Saved to: path for cached
Markdown and surfaces the winning source label (Europe PMC XML, PMC HTML,
etc.) in Markdown and JSON provenance. For XML/JATS winners, the saved Markdown
keeps section text, tables, references, figure captions, supplementary-material
metadata, and explicit markers for complex merged-cell tables that are not yet
flattened. JSON fulltext responses also include full_text_manifest, an
additive artifact manifest with the normalized source family (jats_xml,
pmc_html, or pdf), provider label/source, concrete source identifier,
quality flags, known license/reuse state, and provenance facts such as
open-access and explicit PDF fallback status. JSON fulltext also reports
not_included counts and points to the OA package asset manifest when figure
images, supplementary files, or complex tables are not inlined.
Requested full-text JSON adds full_text_coverage. Its final coverage is
full_text, abstract_only, metadata_only, none, or unavailable and its
ordered attempts list gives each eligible provider's source kind, coverage,
outcome, cache state, and bounded reason. Attempts are sanitized: they do not
contain source bodies, errors, URLs, credentials, parser details, or local
paths. Ordinary article cards omit this object.
JSON always exposes section_outcomes.fulltext: base cards are
not_requested, successful retrieval is data, an all-healthy ladder with no
winner is empty, and a ladder with any failed eligible source and no winner is
unavailable. Structural coverage is independent, so an abstract or metadata
partial remains visible even when another source failed. _meta.section_sources
mirrors the section outcome and provider list. Markdown uses the same combined
state: it distinguishes confirmed absence, partial metadata without an article
body, and incomplete retrieval without listing provider attempts.
Article assets:
For a task-oriented walkthrough of listing, downloading, and interpreting these files, see Supplementary Materials.
get article <id> assets returns a JSON-only provider-labelled manifest. BioMCP
merges PMC OA's versioned S3 metadata-declared media objects, a validated Europe PMC supplementary ZIP,
recognized JATS/PMC HTML supplement links, and eligible Figshare/AACR metadata. PMC OA no longer uses
the retired FTP/archive route, whose legacy files are removed in August 2026. Figshare
manifests can include same-paper sibling records discovered by DOI/title, while
excluding records that do not match the paper. Linked provider URLs stay internal;
coverage reports named files that are absent, denied, unsupported, unavailable, or gated.
The manifest's source_attempts distinguishes data, degraded data, healthy
absence, provider failure, and timeout. Optional-provider work has an overall
budget, so a slow optional source cannot withhold assets already returned by a
successful source. Normal calls reuse the resolved manifest for five minutes,
including paging continuations; --no-cache bypasses that reuse.
PMC's pmc_proof_of_work outcome means PMC returned a proof-of-work challenge
instead of the named file; it has no asset_key or handle. A linked binary
received as HTML or XHTML is rejected rather than exposed as asset bytes.
get article <id> asset <asset-key>
returns the selected member as raw bytes with no conversion; downstream tools
parse CSV, XLSX, DOC, PDF, or images. Manifest handles remain BioMCP commands,
not provider download URLs. Europe PMC reuse remains unknown unless article
metadata or a retained PMC OA manifest supplies a license; retained licenses
name PMC OA as their source. When named coverage exists but no bytes are
retrievable, the manifest remains available with assets: [] and a specific
per-file outcome. If no route names an asset, a healthy all-source miss is
not_found; source failures remain source_unavailable. A missing or unknown
asset key is not_found. Figshare supplement PDFs and tables remain assets,
not full-text article substitutes. Full-text downloads are managed private
state (0600 files on Unix); routine managed-state maintenance also repairs
older permissive files.
Opt in to the final PDF rung only when you want the last-resort open-access PDF path after XML and PMC HTML fail to provide an article body (including when they provide only an abstract or metadata):
With --pdf, BioMCP can use a Semantic Scholar open-access PDF URL from an
explicitly allowed Semantic Scholar/CDN HTTPS origin as the final fallback and
labels the winner as Semantic Scholar PDF. A successful lookup with no PDF is
a healthy absence; a failed lookup contributes unavailable state only when PDF
fallback was requested. Other provider-returned origins are rejected before
contact. --pdf is only valid with the fulltext section;
biomcp get article 22663011 --pdf is rejected instead of silently doing
nothing.
Indexing section:
This opt-in section fetches PubMed citation XML and keeps each author's name,
optional ORCID, and source-associated affiliations together. MeSH descriptors
and qualifiers retain their UIs and independent major-topic flags. The explicit
available or unavailable status distinguishes a complete empty author/MeSH
list from metadata BioMCP could not retrieve. BioMCP accepts PubMed's normal
external DOCTYPE without downloading the DTD, while retaining the 8 MiB body
limit and a 100,000-node XML limit. When indexing is unavailable, JSON and
Markdown include a sanitized failure code and static message (for example,
rate_limited, parse_error, or timeout) without upstream bodies, URLs,
credentials, or parser details. The base article still succeeds. Ordinary detail,
search, and batch do not make this extra request; get article <id> all includes
it.
Annotation section:
Semantic Scholar TLDR section:
The attempted Semantic Scholar enrichment is recorded in
section_outcomes.tldr as data, empty, or unavailable; its
_meta.section_sources entry projects the same state and credits Semantic
Scholar only after a successful response.
Helper commands¶
biomcp article entities 30738221 # annotated entities with identifiers and follow-up get commands
biomcp --json article entities 30738221 --full # same rows with passage positions
biomcp batch article 22663011,24200969 --mode compact # compact summary cards
biomcp batch article 22663011,24200969 # detail records (default)
biomcp article citations 22663011 --limit 3 --offset 0 # one Semantic Scholar citation page
biomcp article references 22663011 --limit 3 --offset 0 # one Semantic Scholar reference page
biomcp article recommendations 22663011 --limit 3 # Semantic Scholar related papers
biomcp article citation-evidence 22663011 10.1038/nature10725 # passage for one directed citation
batch article --mode compact works without S2_API_KEY and returns the
standard {summary, items} batch envelope in request order. Each compact card
echoes the original requested_id, keeps its
resolved PMID/PMCID/DOI fields, and carries the same full authors, returned
author_count, author_completeness, and author_source contract as detail.
Markdown cards show the source-ordered names and an explicit status, including
when no author list was supplied. When Semantic Scholar data is available, the
batch helper can add optional TLDR and citation metadata without changing
authorship. S2_API_KEY makes that enrichment authenticated and more reliable.
Use batch article --mode compact as the default follow-up after search article when you
already have several shortlisted PMIDs or DOIs.
The Semantic Scholar graph helpers also work without S2_API_KEY, but they use
the shared pool and can fail fast on HTTP 429 with guidance to set the key for
a dedicated rate limit. Citations usually work broadly; references and
recommendations can be sparse or empty for paywalled papers because of
publisher elision in the Semantic Scholar graph. In JSON, citation/reference
responses always carry edges, and recommendation responses always carry
recommendations, including [] on successful emptiness and parsed structured
errors. On errors, the accompanying error and nonzero exit status remain
mandatory, so an empty array must not be interpreted as a biomedical negative.
Each citation/reference invocation returns exactly one requested --offset
page. Follow only the provider-authored command in _meta.next_commands or the
Markdown Next: line; coverage_status: exhausted means that response omitted
a continuation, not that BioMCP calculated a total.
article citation-evidence <citing-id> <cited-id> returns bounded source text
for one directed citation pair: Semantic Scholar context by default, and when
the edge carries no context, the paragraphs of the citing paper's open
Europe PMC JATS full text whose unambiguous bibliographic markers link to the
cited reference. Pass --fulltext to force the JATS path even when provider
context exists. When no passage is openly available, OpenCitations confirms
whether the directed edge exists and the response reports that confirmed edge
without a passage. The response is a closed six-status evidence object
(context_from_provider, context_from_fulltext, fulltext_unavailable,
reference_confirmed_without_passage, reference_unresolved,
citation_marker_unlinked) with per-passage locators and evidence URLs. It
retrieves evidence only: BioMCP does not summarize the passage or interpret
how the cited work was used.
Caching behavior¶
Downloaded content is stored in the BioMCP cache directory. This avoids repeated large payload downloads during iterative workflows.
JSON mode¶
biomcp --json get article 22663011
biomcp --json search article -g BRAF --limit 3
biomcp --json search article -k "Oncotype DX review" --session lit-review-1 --limit 5
biomcp --json batch article 22663011,24200969 --mode compact
JSON article responses include _meta.next_commands and _meta.section_sources,
so article workflows can promote the next likely pivots and preserve section
provenance without scraping markdown. For get article <id> fulltext --json,
full_text_manifest.quality reports whether saved Markdown has sections,
tables, references, fulltext signal, and fulltext entity annotations;
full_text_manifest.reuse separates known licenses from an explicit unknown
license/reuse warning; and full_text_manifest.provenance carries available
open-access, retraction, package, and PDF-fallback facts. JSON search article
responses also echo query, sort, semantic_scholar_enabled, and row-level ranking/provenance
metadata. In relevance mode, ranking metadata now includes the effective mode
plus normalized lexical, citation, and position components; semantic-aware
rows expose ranking.semantic_score as the LitSense2-derived signal and use
0 when LitSense2 did not match. Hybrid rows also include the composite
score. Keyword-only article searches with an exact gene, drug, or disease
label/alias match may include _meta.suggestions[] objects with command,
reason, and sections; same-session keyword loop-breaker suggestions include
command and reason and omit sections. _meta.next_commands remains the
executable string command list. JSON batch article responses use the standard
batch envelope so callers can map settled results back to the original input order.
The compatibility spelling biomcp article batch <id1> <id2> ... remains
available with its existing compact-card output; prefer the canonical command
in new scripts.
Practical tips¶
- Start with narrow
--limitvalues. - Add a disease term when gene-only search is too broad.
- Use section requests to avoid oversized responses.
- Use
biomcp get article <id> indexingfor PubMed author-affiliation and MeSH indexing metadata. - Use
biomcp get article <id> tldrwhen you want only the optional Semantic Scholar section.