File Description: go-stats

From GO Wiki
Jump to navigation Jump to search

Usage

Primary stats file computed.

Input data

Annotation stats are obtained by querying the GOlr (GO Solr instance).

Format(s)

json

File description

The go-stats file contains the following information:

release_date

  • release_date: Obtained from release/metadata/release-date.json or snapshot/metadata/release-date.json.

ontology

  • valid_terms: Total number of valid terms (non-obsolete) in the ontology.
  • obsolete_terms: Total number of terms with obsolete status (ie, term_ids for which the is_obsolete field is true in the go.obo file) (this excludes merges).
  • merged_terms: Total number of merged terms (calculated by counting the term_ids for which the field is_obsolete is true in the go.obo file, and that also are are as alt_ids of a valid term).
  • biological_process_terms: Total number of valid terms for the biological_process aspect.
  • molecular_function_terms: Total number of valid terms for the molecular_function aspect.
  • cellular_component_terms: Total number of valid terms for the cellular_component aspect.
  • meta_statements: Total number of identifiers, alternative identifiers, namespace, term label, comments, synonyms, definitions, subsets, for each valid term.
  • cross_references: Total number of cross_references, from the xref field of the go.obo file.
  • terms_relations: Total number of relations; the count of all relations, using the fields is_a, intersection_of and relationship of the go.obo file.
  • changes_created_terms: Number of created terms since the previous release.
  • changes_valid_terms: Number of valid terms since the previous release.
  • changes_obsolete_terms: Number of terms obsoleted since the previous release.
  • changes_merged_terms: Number of created merged since the previous release.
  • changes_biological_process_terms: Changes in the number of BP terms.
  • changes_molecular_function_terms": Changes in the number of MF terms.
  • changes_cellular_component_terms":Changes in the number of CC terms.

annotations

  • total: The total number of annotations.
  • by_aspect: P, F, C.
  • by_bioentity_type:
  • by_qualifier |by_qualifier]]: contributes_to, colocalizes_with, NOT
  • by_taxon: Number of annotations for each of the annotated species in the database.
  • by_evidence
  • by_model_organism: For each species, the number of annotations are shown:
  • by_group: Number of annotation for each contributing group, obtained using the assigned_by field of each input file.

taxa

  • taxa: Number of species with annotations.
  • taxa_filtered: Number of species with at least 1,000 annotations.

bioentities

references

  • all
    • total: Total number of distinct annotated references (includes PMIDs, GO_REFs, DOIs, internal IDs for Model Organism Databases and Reactome (note that for papers with both a PMID and an internal reference ID, the paper is counted twice).
    • by_filtered_taxon: Total number of annotated references by species.
    • by_group: Total number of annotated references for each contributing group, obtained using the assigned_by field.
  • pmids
    • total: Total number of annotated PMIDs.
    • by_filtered_taxon: Total number of annotated PMIDs by species.
    • by_group: Total number of annotated PMIDs for each contributing group, obtained using the assigned_by field.

Direct access to files

snapshot

http://snapshot.geneontology.org/release_stats/go-stats.json

current

http://current.geneontology.org/release_stats/go-stats.json

Review Status

Last reviewed: October 24, 2019