Reference Genome Database Requirements Discussion 2007 (Retired)

From GO Wiki
Jump to navigation Jump to search

This is the place to discuss features and requirements for the Reference Genome Database being designed to replace the Google Spreadsheet system currently in use.

(here's one to get us started --chris):

Ensures consistent use of identifiers

Identifiers must unambiguously identify a single entry in a database.

Identifiers should conform to the following syntax:

 DBAuthority : LocalID

DBAuthority should be in the GO xrefs metadata list:

E.g.

 FB:FBgn0000001

Curators should not be expected to memorise identifiers, so a data entry system should allow them to enter symbols etc and have this resolved as an ID eg using some automatic lookup mechanism

Should allow loading of MOD reports

The database should allow MODs to submit their metrics via a tab-delimited file that can be automatically downloaded from their ftp site. The file should contain columns for the Reference gene, the organism's ortholog/orthologs, date genes have been completed and reference counts for total number of papers associated with a gene etc.

In the future we may want to add capability to determine when genes that have been completed but have new references associated with them.

References should be compiled in a central location, so once a paper is curated, it is somehow flagged that it has been done.

ZFIN does this for our own local publication database. It is very handy since one paper typically touches on several genes. Completing GO for a paper means it doesn't need to be looked at again if it is associated with another gene on the ref genome list. This would be useful for groups that don't have such a mechanism in-house. MODS could provide a tab delimited file like geneID | PUB ID | OMIM ID | pub status [curated/not curated]. If the Ref. Genome interface had this data, it would also be possible to show a report of which pubs still need to be examined for a specific ortholog. One possible complication is that a single pub may deal with the genes from multiple species. Would need to track curation of this pub separately for each species I think. -Doug

We need to decide whether we will allow individual users to modify the database a record at a time or whether the database should only be populated with files from each MOD.

I think both methods would be useful (Doug)

Should track that no ortholog was found

There should be a mechanism for indicating that a curator has looked and no ortholog could be located as of a certain date. For genomes that are not yet completely sequenced, we will want to revisit these when a new genome build is released. It would be nice to have a free text note field associated as well so we can leave notes regarding the analysis that was performed. -Doug

Should provide reports to focus curation effort

The interface should provide reports that will help focus curator effort. One example might be to provide a facility to search the data for species-specific orthologs where curation is not 'comprehensive'...these are the genes we should be working on. Another example might be to provide a report of genes where no ortholog was determined yet. It would be nice to be able to alter the sort order of the results in such reports by the following parameters: date Human Gene was added to the Ref. Genome set, ortholog ID, OMIM ID, most papers associated, difference between papers associated with a gene and papers read for GO for that gene (biggest difference to smallest)...maybe others... It would also be good to have a report generated by a query on OMIM ID that shows which species have 'comprehensive' annotation done for their ortholog(s)-Doug

Should record that orthology is 'comprehensive' as of a certain date

Curators should be able to mark that curation of an ortholog is 'comprehensive' as of a certain date. It would be good to be able to generate a report to look for cases where the 'comprehensive' curation date is getting old. These may need to be reviewed and updated. -Doug

Should allow a 1:many relation between Human gene and MOD ortholog

Perhaps this is self-evident, but it mucks up reports and things pretty nicely...

Should record orthology determination method

If the group decides on a single method we can all adhere to for ortholog determination, then this may not be required. I think this is unlikely however, in which case it may be useful to record HOW orthology was determined. Perhaps a set of checkboxes beside the various methods can be used to record what tools/methods supported the orthology call. These would include YOGY, TreeFam, InParanoid, etc. as well as manual BLAST and synteny analysis, and perhaps 'established from literature' for cases where orthology was already recorded by a MOD from a published paper. If we really want to be careful, it might be good to record which build of the databases behind the tools was being used as well..ie which Sanger build, which InParanoid version, etc.? -Doug

Only one index page

Ideally one main page with options for the pages to visit or output required. Also all pages to have links back to index and too any pages the page is linked from, ie if you edit a page you can't use the back button and don't want to have to keep looking for a bookmarked page (Ruth)

Easy data input

The administrators and curators need an easy way to add 20 genes at a time to a table. The curators need an easy way to view the genes done/to do and to edit as they go.

Would it be possible to that when the admin create a new gene record automatically 12 pages are created for a gene in all of the other species?

Then either each gene page could have a table listing a synopsis of the annotation so far achieved in all other species; or just the human gene page would have this data: eg 1 row (unless paralogs) per species column headings: for each species: annotator to contact for gene discussions, in progress/date completed, annotations added.

as well as fields with dropdown choices "paralog, ortholog, ancestral gene" fields for metric data, dropdown choice "annotations added: yes/no" curator assigned to gene etc

Obviously also need option of duplicating page in cases where there are paralogs, duplication would ensure link to initial gene page is maintained. (Ruth)

Administrator output for easy data retrieval

Would it be possible to select different output options, ie html or excel?

The advantage of excel is that people can manipulate the data as they wish, unless a variety of outputs, eg graphs, data collation can be included in the outputs of this tool.

I would suggest that the administrators would appreciate an output table which is similar to the original google spreadsheet. With each human gene listed in separate rows (in cases where there are paralogs there will obviously be multiple rows/gene), and the accession number and the metrics data and the date completed for human, and all other species listed in columns.

However, I don't think they will want to view the table as a whole every time they look at it. Especially in a couple of years time when there are 500 genes on the list.

Therefore could there be drop down options: eg having selected "metrics table" and then "edit" or "view", then for view have options "excel" or "html" then next options are: "all data", "by date added to table", "only genes comprehensively annotated in all species", "newest genes", "genes not yet annotated" this would I guess lead to the further option of dates, "2006", "Aug06", "Sept06"..., alphabetically. Perhaps the choices should be decided once people work out what data they want. (Ruth)

Curator output for easy view of data

I think curators would appreciate options (similar to above) for viewing a "spreadsheet". I don't think they will want to look at the whole table every time.

Therefore could there be drop down options: eg having selected "curator table" and then "edit" or "view", (for view have options "excel" or "html") then next options: choice of "species", human, mouse etc; then choice of "genes", "all genes", "by date added to table", "only genes comprehensively annotated in all species", "newest genes", "genes not yet annotated" "genes not yet assigned to curator", "genes assigned to curator...Ruth" this would I guess lead to the further option of dates, "2006", "Aug06", "Sept06"..., alphabetically. Perhaps the choices should be decided once people work out what data they want.

Ideally the species specific spreadsheets would contain all the species specific data available in the individual gene records so that people could edit the spreadsheet rather than use the gene records if they wanted to. (Ruth)

Comments on prototypes

Prototypes for the curation tool can be found here: http://rails-dev.bioinformatics.northwestern.edu:24000/curator

http://rails-dev.bioinformatics.northwestern.edu:24000/admin

Please add your comments & suggestions here

Susan

  1. Curator Central - adding orthologs:

To test this, I made (and destroyed) a new Drosophila ortholog (of HK001?) called Stuff - this worked fine. I then looked in Admin Central and was surprised to see that Stuff was present in the 'listing of target genes' - shouldn't this have disappeared for this listing as I had already 'destroyed' the # ortholog in the other table? I was also suprised to see my new Dros ortholog in this table of target genes because I had expected 'target genes' to be the original list of human genes. Are we going to distinguish between 'target genes set' (jn human) and 'target genes identified' (orthologs in other species)? If not then we need to be able to view by species. - Susan

A related issue applies to the table for 'Listing Curation Status'. I assume we will view 'Listing Curation Status' for one species at a time? For those of us not curating human papers, I'd like to see 2 columns on the left - one for the human gene symbol and one with the ortholog symbol. The current spreadsheets include many symbols for the human gene - I assume the human symbol used as in these summary sheets will be something standard like such as the valid HGNC symbol? - Susan

Should 'Complete curation' be entered as a date rather true/false - Susan


Val

  1. Will there be a page like this


http://rails-dev.bioinformatics.northwestern.edu:24000/track_curation/list


for each ortholog, or for each organism?

  • Sohel's answer

There will be. I'm guessing for each organism to mirror the current spreadsheet , but it might better by ortholog so that we could easily look at the situation in other species as we are annotating.

  • There would be a separate view for each model organism. Only the curators

of that organism would be able to see/update their respective curation status. But again this is open for discussion. Additionally, we could have the Ortholog view, such that curators can look at the orthologs in other species.


  1. There does not appear to be a column for the annotated species ID? (or a species coulumn, if this is by ortholog)

Sohel's answer

  • This can be added.
  1. Would it be possible to display a couple of 'real' reference genome entries so we can see how this would look/link together for the different species?

Sohel answer

  • I think it is a good idea. We will work on getting some real data
  1. If this was by ortholog perhaps we could also include direct x-links to

i) Marys graphs ii) Treefam iii) YOGY iv) the uniprot and hugo entry for the human protein at the top or bottom of the page. Would this be useful ? Sohel answer

  • We are planning to put these links on the "Add Ortholog" page

Ruth However before the meeting please could a few modifications be made to the current pages so that it is easier to get forwards and backwards to pages. Would it be possible for each page could have a back option and a "index" option or something along those lines.


Also the headings <http://rails-dev.bioinformatics.northwestern.edu:24000/feature/list#>Target Gene/Gene product id/symbol and <http://rails-dev.bioinformatics.northwestern.edu:24000/feature/list#>Gene/Gene product are confusing. From what I can see I think it is more helpful to be choosing an otholog by looking at the gene name and the protein ID. However if only one on view then I guess it should be a unique identifier. Therefore the <http://rails-dev.bioinformatics.northwestern.edu:24000/feature/list#>Target Gene/Gene product id/symbol heading should be protein accession/internal ID and the <http://rails-dev.bioinformatics.northwestern.edu:24000/feature/list#>Gene/Gene product heading should be gene symbol. But maybe I am doing this wrong and that is why it keeps crashing!



your name comments