Skip to content

Variant

Use variant commands for compact annotation, source-backed interpretation context, and optional predictive and population-genetics sections.

Accepted variant identifiers

BioMCP supports multiple input forms:

  • rsID: rs113488022
  • ClinVar VariationID: 577152
  • HGVS genomic: chr7:g.140453136A>T
  • exact-copy repeat: chr19:g.11106928AAG[1]
  • range deletion: chr2:g.47641567_47641569del
  • sequence-qualified deletion: chr19:g.11106928delAAG
  • duplication: chr19:g.11106928dup
  • insertion: chr2:g.47641567_47641568insA
  • inversion: chr2:g.47641567_47641569inv
  • delins: chr2:g.47641567_47641569delinsA
  • build-qualified genomic: GRCh38:chr7:g.140753336A>T or hg19:chr7:g.140453136A>T
  • versioned RefSeq HGVS: NC_000010.11:g.87925512G>A
  • VCF-like: chr10:87925512:G:A
  • SPDI: NC_000010.11:87925511:G:A
  • versioned RefSeq deletion: NC_000010.11:g.87925512del
  • transcript coding HGVS: NM_000249.4:c.678-14_678-3del (intronic offsets resolve through ClinVar's coding aliases when normalization services refuse)
  • gene-protein form: BRAF V600E, BRAF p.Val600Glu, EGFR E746_A750del

ClinVar-style names that keep the gene in parentheses, such as NM_000249.4(MLH1):c.678-14_678-3del, are refused with the colon working form printed in the message; BioMCP never half-parses them.

These exact formats are accepted by biomcp get variant and the exact-ID helper commands. Use --assembly hg19 or --assembly hg38 to declare the build for a chromosome-prefixed coordinate; build-qualified and versioned RefSeq inputs already establish it. For a bare coordinate, BioMCP checks both GRCh37 and GRCh38, preferring GRCh38. Set BIOMCP_DEFAULT_ASSEMBLY=grch37 (or hg19) to restore GRCh37 preference; an explicit --assembly always wins. If the same spelling identifies different records, BioMCP returns the preferred record and reports the other identity in build_candidates JSON and a Markdown warning.

A gene-protein form names a protein change, and the same change can sit on several genomic variants; provider alias lists even carry other isoforms' spellings on different variants. get variant counts a hit as a match only when the transcript BioMCP headlines for it spells the requested change, and that headline prefers the MANE transcript the response marks (ClinVar names variants on MANE Select) over the first NM_ annotation. When the resolved change does not spell the request, the answer checks the requested reference residue against the gene's canonical (MANE Select) protein sequence: a matching residue means the request's own numbering is valid on MANE, and if no record names the change there, get variant refuses with the only alias match as a candidate; otherwise the answer carries a numbering note naming the transcript and spelling that matched. When exactly one matching variant carries the ClinVar record for that change, get variant resolves to it; when several variants match and none or several carry ClinVar records, it refuses with the candidates and a working input form instead of guessing. Retry with one of the listed genomic HGVS, rsID, ClinVar VariationID, or transcript-qualified HGVS spellings.

ClinGen Allele Registry normalization

Use CAR for a source-provided CAid and bounded alias collections for versioned RefSeq transcript coding (NM_...:c.) or genomic (NC_...:g.) HGVS values. It is a read-only lookup: BioMCP does not infer equivalence, register alleles, or liftover.

biomcp --json variant normalize car 'NM_000546.6:c.215C>G'
biomcp --json variant normalize car --input car-hgvs.json

The batch file is a bare JSON array of 1–50 values and may also be read from stdin with --input -. --input is available only with car and requires --json.

ClinGen ERepo expert assertions

Use a ClinGen Allele identifier (CAid) to retrieve source-faithful expert-panel assertion summaries. The result keeps met and unmet source lists separate and does not infer a criterion strength from defaults or comments.

ClinGen ERepo contains germline variant interpretations. For somatic tumor questions, use CIViC: get gene civic or get variant civic.

biomcp variant erepo CA015543
biomcp --json variant erepo CA015543 --detail
biomcp --json variant erepo --input caids.json
biomcp --json variant erepo --gene PTEN --limit 25 --offset 0

Batch input accepts 1–50 CAids and is summary-only. Detail fetches one assertion; when a summary has multiple assertions, supply --assertion <UUID>, and use --version <exact-docVersion> only with --detail.

Gene mode is mutually exclusive with CAID and batch/detail modes. It returns a bounded compact page, requests one extra provider row to derive has_more, and reports total: null because ERepo does not provide an exact total. At most three HGVS strings are shown per assertion; the response retains the full count and names any oversized preview fields it omitted.

Search variants

By gene and protein change:

biomcp search variant -g BRAF --hgvsp V600E --limit 5
biomcp search variant -g BRAF --hgvsp p.Val600Glu --limit 5
biomcp search variant BRAF p.Val600Glu --limit 5

By residue alias shorthand:

biomcp search variant "PTPN22 620W" --limit 5

By protein shorthand when gene context is already supplied:

biomcp search variant -g PTPN22 R620W --limit 5

Standalone protein shorthand like R620W returns variant-specific recovery guidance instead of falling back to gene or condition discovery.

Protein searches using --hgvsp or positional GENE CHANGE, plus coding-HGVS and rsID searches, are exact identity routes. BioMCP keeps the supplied spelling, normalizes aliases only for comparison, and excludes MyVariant candidates whose source gene or allele facts contradict the request. Exact JSON includes requested_variant and a resolution status: resolved for one source-proved identity, ambiguous when source evidence is incomplete or multiple identities remain, and unresolved when exhaustive candidates are absent or contradictory. Retained rows expose the provider's complete source_identity arrays and the source-derived matched_alias that proved the match.

Exact-search rows also include transcript_annotations_complete and a bounded transcript_annotations array. Each object preserves one intact MyVariant.info snpeff.ann gene/transcript/coding/protein tuple and marks it as displayed, matched, both, or neither. The displayed role follows BioMCP's existing SnpEff display selection; the matched role is present only when that same SnpEff object satisfies every submitted gene and transcript-specific selector. The independent dbNSFP arrays remain the evidence used to retain an exact result, so matched_alias can be present when no SnpEff tuple is marked matched. These annotations explain source representations; they do not choose a preferred transcript, establish equivalence, or convey clinical meaning. Broad searches omit both fields. An empty array with false means malformed or over-budget SnpEff data was isolated while the otherwise usable result was retained.

Search output reports that exact result as Variant identity in Markdown and as resolution in JSON. A separate Filter evaluation block (the filter_evaluation JSON object) marks each submitted filter as evaluated or unavailable. evaluated means the filter participated in an interpretable provider query; it does not mean that the filter value names a record. This is why an exact identity can be unresolved while its gene and protein filters are both evaluated. An unavailable gene may appear beside another evaluated filter when provider evidence cannot support a reliable negative gene result. Broad searches omit requested_variant and resolution but keep their filter evaluation object, including an empty object when no filters were submitted.

Gene-only, residue-alias, significance, consequence, score, population, and other discovery filters remain broad searches and omit strict identity metadata.

By significance:

biomcp search variant -g BRCA1 --significance pathogenic --limit 5

With population and score filters:

biomcp search variant -g BRCA1 --max-frequency 0.01 --min-cadd 20 --limit 5

Consequence, ClinVar review-status, and field-presence filters can be combined:

biomcp search variant -g BRAF --consequence missense_variant --limit 5
biomcp search variant -g BRCA1 --review-status 2 --limit 5
biomcp search variant -g BRAF --has revel --limit 5
biomcp search variant -g BRAF --missing revel --limit 5

--has and --missing accept cadd, revel, gerp, clinvar, gnomad, dbsnp, snpeff, civic, and cosmic. Unsupported consequence, review-status, or field values exit with a typed invalid_argument error rather than reporting an empty successful search. The same is true for contradictory predicates: a field cannot be requested with both --has and --missing, and a field-specific filter cannot be combined with --missing for that field. These errors occur before provider contact.

CADD and GERP thresholds must be finite. Results filtered with --gerp-min include the qualifying gerp score in each row.

Get a variant record

biomcp get variant rs113488022
biomcp get variant "chr7:g.140453136A>T"
biomcp get variant "BRAF V600E"
biomcp get variant "BRAF p.Val600Glu"

The default output favors concise, clinically relevant context first. When BioMCP can derive a compact literature-facing alias, variant search rows and detail cards also show Legacy Name, and JSON output includes legacy_name.

Shorthand such as PTPN22 620W or R620W is not treated as an exact variant ID. Use biomcp search variant for those inputs.

Normalize transcript HGVS

Use the explicit normalization proxy when you already have a transcript HGVS string and want source-labelled output from Mutalyzer and VariantValidator:

biomcp variant normalize all NM_000248.3:c.135del
biomcp variant normalize all 'NM_004448.2:c.829G>T'
biomcp variant normalize mutalyzer NM_000248.3:c.135del
biomcp variant normalize variantvalidator 'NM_004448.2:c.829G>T'

JSON output always writes a parseable object on exit 0. It preserves the submitted input, an aggregate status, a results list of normalized forms, a message, one result per service with each service status, source-returned transcript/normalized/genomic/protein fields, warnings such as VariantValidator TranscriptVersionWarning, and _meta.next_commands. When no provider returns a normalized form, status is no_result and results is an empty list. Markdown VariantValidator genomic descriptions are labeled as GRCh38 because BioMCP calls VariantValidator's GRCh38 normalization endpoint.

This helper does not parse messy report prose, does not choose, select, guess, or infer transcripts, and does not classify variants or provide clinical interpretation/clinical meaning.

Request variant sections

The default variant card includes a one-line therapeutic-evidence pointer. When cached MyVariant CIViC evidence is already in the fetched variant payload, the line reports the cached predictive-item count and points to get variant <id> civic; otherwise it prints the bare get variant <id> civic next-command. The default card does not run the live CIViC GraphQL section call.

Prediction section:

biomcp get variant "BRAF V600E" predict

ClinVar-focused section:

biomcp get variant rs113488022 clinvar

ClinVar JSON also exposes top_disease when condition aggregation is available, reusing the highest-ranked ClinVar condition row already shown in the section. The runtime See also: block is significance-aware: pathogenic or likely pathogenic variants keep the gene and drug-target pivots near the top, while VUS / uncertain-significance variants add a literature search anchored by the variant alias and any available disease context before the generic drug-target fallback.

Population section:

biomcp get variant "GRCh38:chr7:g.140753336A>T" population

Population Markdown shows compact exome and genome frequencies, the highest observed population-row frequency with its allele count and allele number, grpmax FAF95, and quality status. Allele number is the AF denominator: the number of alleles with a defined genotype call at that site. Use the full-table route when you need every population row:

biomcp get variant "GRCh38:chr7:g.140753336A>T" population-details

Population JSON exposes a direct gnomad_r4 result with separate exome and genome objects. Each object preserves raw numeric allele frequencies and counts by ancestry, quality flags, and grpmax FAF95 when available. The result distinguishes a missing trustworthy GRCh38 coordinate, a variant absent from gnomAD v4, and a provider failure. gnomAD excludes bottlenecked genetic ancestry groups when selecting grpmax FAF; BioMCP repeats that caveat in the result.

CIViC section:

biomcp get variant "BRAF V600E" civic

The CIViC section includes cached MyVariant rows plus live GraphQL context when available. It also carries a currency caveat and points to literature and drug cross-check commands because curated CIViC evidence can lag current standard of care.

GWAS section (trait associations from GWAS Catalog):

biomcp get variant rs7903146 gwas

GWAS JSON exposes supporting_pmids as an ordered, deduplicated array when the section loads successfully. null means either the GWAS section was not loaded or GWAS Catalog was temporarily unavailable. In the unavailable case, JSON also includes gwas_unavailable_reason with a source-specific message.

Predictions (aggregated prediction scores):

biomcp get variant "BRAF V600E" predictions

MyVariant-backed prediction entries can include REVEL, AlphaMissense, ClinPred, SIFT, MetaRNN, BayesDel add-AF, and BayesDel no-AF. The two BayesDel flavors remain separate source scores; BioMCP does not apply a clinical threshold or classify pathogenicity from either score.

Conservation (GERP, phyloP):

biomcp get variant rs113488022 conservation

COSMIC (somatic mutation data):

biomcp get variant "BRAF V600E" cosmic

CGI (Cancer Genome Interpreter annotations):

biomcp get variant "BRAF V600E" cgi

cBioPortal (frequency data):

biomcp get variant "BRAF V600E" cbioportal

Cancerhotspots.org recurrence counts are loaded by all for exact gene/protein queries such as BRAF V600E. JSON includes cancerhotspots.source, matched_transcript, position_count (residue-level tumorCount), and same_aa_count (the exact alternate amino acid count). If the lookup succeeds but the residue/change is not a cancerhotspots hotspot, those count/provenance fields are present as JSON null under the source-labelled object; upstream unavailability omits the object instead of emitting zeros.

All supported sections:

biomcp get variant rs113488022 all

all includes broad source-backed annotation sections such as ClinVar, population, conservation, expanded MyVariant prediction scores, cBioPortal, Cancerhotspots.org, CIViC, CGI, COSMIC, and GWAS. It does not run the AlphaGenome predict section because that path requires ALPHAGENOME_API_KEY; request predict explicitly when you want AlphaGenome output.

Helper commands

biomcp variant trials "BRAF V600E"     # search trials mentioning this mutation
biomcp variant articles "BRAF V600E"   # union exact article routes for this variant
biomcp variant structure "BRAF V600E"  # residue/domain/PDB/AlphaFold/hotspot context
biomcp variant oncokb "BRAF V600E"     # OncoKB lookup (requires ONCOKB_TOKEN)

variant structure is an opt-in, network-backed helper. JSON includes the selected residue, matched HGVSp aliases, other MyVariant/dbNSFP positions, overlapping InterPro domain ranges, typed UniProt PDB rows, an AlphaFold URL, Cancerhotspots recurrence, top-level lookup_outcomes, warnings, and _meta.next_commands. lookup_outcomes.domains and lookup_outcomes.cancerhotspots distinguish data, healthy empty, local inapplicable, and temporary unavailable results. Cancer Hotspots recurrence is null when that lookup is inapplicable or unavailable; its successful data and empty object shape is unchanged. This helper does not add structure data to default get variant output.

variant articles first performs one strict variant resolution. Its default --strategy union plans provider-specific strict requests before retaining the ordinary discovery federation and source-backed PubMed citations. BioMCP merges duplicate papers with route/source request provenance, ranks the union, and applies --offset and --limit once. provenance.query_aliases records only the aliases sent to retrieve a row; it does not claim that an article contains or verifies an alias. --debug-plan exposes each provider request, route, query alias, exact query, and template version. It also includes a bounded, versioned candidate-route trace: receipt, union/dedup survival, rank, identity-verification disposition, and pagination visibility for each retained candidate-route observation. Candidate trace v2 reports observed_total, dropped, and computed bounded; observed_total - dropped always equals the length of candidates. It retains visible, deduplicated receipts first, then fills remaining capacity in observation order. Use --strategy annotation or --strategy lexical to diagnose one exact route. Ambiguous or unresolved input still runs strict literal requests unless provider validation contradicts the request; discovery rows remain labeled best_effort_free_text.

JSON repeats requested_variant on every row and reports resolution, complete, truncated, full pagination state, and per-route source_status. An incomplete provider route keeps available rows but sets complete: false, truncated: true, and pagination.total: null. A route stopped before BioMCP makes a provider request is reported as source-neutral internal / not_attempted, rather than attributing an outage to an uncalled provider.

--verify-identity adds captured-evidence identity facts without filtering the retrieved pool. --confirmed-only requires it and filters before ranking and pagination. Query aliases remain retrieval provenance, never observed aliases. Identity observations retain source, section, locator, linked gene, observed alias, typed provider_linkage, and canonical captured-content hash. Confirmation uses returned-PMID typed PubTator Gene/Variant facts with matching exact HGVS and CorrespondingGene/NCBI Gene-ID evidence, or an optional bounded ClinGen LDH annotation that binds the applicable CAid, requested gene, known PMCID, and one exact text/table selector. LDH only observes already retrieved PMCID candidates: missing coverage, malformed data, and outages never remove a candidate or become negative evidence. --debug-plan records versioned clinically relevant response/content subset hashes, the post-response verifier and provider-template versions, plus artifact identity; those hashes are audit facts, never retrieval-cache keys. Frozen fixture coverage is release proof; the matching live probe is diagnostic only. The typed MCP variant_articles tool exposes equivalent verify_identity and confirmed_only controls and compact identity output.

For an exact gene-plus-protein literature review, run biomcp skill exact-variant-literature. It checks identity before the default union shortlist, compares selected summaries with batch article --mode compact, requests full text and assets only for chosen papers, and expands citations or references only when needed.

For several variants, pass a JSON array of 1-10 structured identities from a file or stdin. Each item can use an rsID, complete genomic HGVS, structured genomic coordinates, gene plus protein change, or coding change plus gene/transcript. Versioned RefSeq accepts either genomic: "NC_...:g...." plus an explicit build, or accession, position, ref, and alt plus the build. Builds are GRCh37 or GRCh38; existing chrN identities remain valid.

[
  {"request_id":"atm","genomic":"NC_000011.10:g.108248927T>G","build":"GRCh38"},
  {"request_id":"atm-components","accession":"NC_000011.10","position":108248927,"ref":"T","alt":"G","build":"GRCh38"}
]

caller_supplied means BioMCP accepted the supplied fields as one caller assertion; it validated syntax but did not establish cross-coordinate equivalence. A unique compatible MyVariant identity instead yields provider_confirmed. resolution.basis is one of those values or null, while provider_validation reports confirmed, not_found, indeterminate, contradictory, or unavailable. Its matched_alias is non-null only for confirmation and contradictory_field only for contradiction; invalid items keep resolution: null.

MyVariant outcome Public behavior
unique confirmation resolved/provider-confirmed; exact routes and source citation
exhaustive no record RefSeq resolved/caller-supplied; exact routes; citation skipped without degradation
indeterminate scan RefSeq resolved/caller-supplied; exact routes; incomplete and unknown total
contradictory facts unresolved/null; no exact routes; optional best_effort_free_text only
unavailable RefSeq resolved/caller-supplied; exact routes; incomplete, truncated, unknown total

For caller-supplied RefSeq, exact aliases start with supplied transcript/coding, gene/coding, and RefSeq genomic forms. When both independently supplied versioned RefSeq transcript/coding and complete versioned genomic/build forms are present, the additive canonical_equivalence sibling records CAR CAid agreement, terminal observations, and response hashes. It never changes resolution, chooses a transcript, or performs liftover. Only confirmed CAR aliases can fill unused retrieval-alias slots, where they remain query provenance. BioMCP performs no liftover, accession-to-chr conversion, strand flip, transcript selection, or inferred coordinate generation.

biomcp --json variant articles --input variants.json --limit 10
cat variants.json | biomcp --json variant articles --input - --debug-plan

The batch response preserves input order under items, uses compact article rows, and supplies parseable article-detail follow-ups in _meta.next_commands. At most two items execute concurrently. Work is bounded to 50 logical calls per valid item and 50 times the item count for the request. With --verify-identity, BioMCP protects verification capacity before discovery and expands that reservation as discovery finds page-eligible candidates, so later discovery cannot consume the work needed to verify them. A discovery route stopped by that bound remains incomplete rather than being presented as exhaustive. --debug-plan is JSON-only and adds normalized aliases, route/provider facts, ranking inputs, budgets, work_allocation accounting for discovery and identity verification, stop state, and the bounded candidate-route trace. Its trace contains only identifiers, route/stage dispositions, and bounded rank positions; it excludes provider URLs, credentials, headers, and raw response bodies. Ordinary single and batch output omit plans. The typed MCP variant_articles tool accepts the same structured item fields directly without server-local paths.

Search GWAS associations

By gene:

biomcp search gwas -g TCF7L2 --limit 10

By trait:

biomcp search gwas --trait "type 2 diabetes" --limit 10

Gene and trait searches each use one bounded GWAS Catalog v2 association request. Supplying both filters returns their rsID intersection, and --p-value is applied afterward. Interval search is not supported. The checked --offset + --limit window must be at most 50; JSON reports followable pages under _meta.pagination and distinguishes provider-budget truncation from true exhaustion.

Optional enrichment

Variant base output may include cBioPortal enrichment when available. OncoKB is accessed explicitly via biomcp variant oncokb "<gene> <variant>" and requires ONCOKB_TOKEN.

Prediction requirements

Prediction sections may require ALPHAGENOME_API_KEY depending on source path. Unsupported command inputs still use explicit validation messages. A valid variant card whose requested lookup lacks a prerequisite remains successful and reports that lookup as inapplicable.

JSON mode

biomcp --json get variant "BRAF V600E"
biomcp --json get variant rs7903146 gwas
biomcp --json search gwas --trait "type 2 diabetes"

JSON records requested predict, cancerhotspots, civic, cbioportal, and gwas source states in section_outcomes; _meta.section_sources projects the same state and successful providers. If genomic coordinates, a gene/protein change, a gene, or an rsID required by one of these lookups is absent, the result is inapplicable: no provider was contacted or credited, and the message names the missing prerequisite. A contacted Cancer Hotspots no-match is empty, while a source failure is unavailable and receives no provider credit.

Practical tips

  • Use search variant first for shorthand or ambiguous inputs.
  • Start with the base card, then add source sections such as clinvar, civic, or population only when needed.
  • Use all when you need a one-shot export for review or downstream comparison.