Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

locateVariants returns genes from both Forward and Reverse strands in PRECEDEID and FOLLOWID

Đang mở
#55 3 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
25/100
Loại issue
Tài liệu
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Đình trệ
Công nghệ
r
Lĩnh vực
bioinformatics

Hướng nghiên cứu

Bắt đầu với tài liệu và phần triển khai của locateVariants() và IntergenicVariants(upstream=1000000, downstream=1000000), sau đó kiểm tra cách PRECEDEID và FOLLOWID được điền và strand được xử lý như thế nào. Được xem là hoàn tất khi cách diễn giải có xét đến strand được ghi lại rõ ràng, hoặc khi hành vi được báo cáo có một regression test tập trung và một bản sửa được thống nhất.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Hello,

I'm looking at an intergenic variant (in the bovine genome) and want to figure the genes relative to which it is downstream or upstream, i.e. relative to the gene's position and its strand. I used locateVariants() with region=IntergenicVariants(upstream=1000000, downstream=1000000) in order to do this. However, some results puzzle me.

Here's what my variant looks like in the gene annotation results:

> all_var_df[rownames(all_var_df) == "AX-106756303", ]

             seqnames    start      end width strand   LOCATION LOCSTART LOCEND QUERYID TXID CDSID GENEID    PRECEDEID     FOLLOWID
AX-106756303        1 34617002 34617002     1      * intergenic       NA     NA       1 <NA>         <NA> ENSBTAG0.... ENSBTAG0....

Here's what PRECEDEID looks like:

lapply(all_var_df[rownames(all_var_df) == "AX-106756303", ]$PRECEDEID, function(X) {mapIds( org.Bt.eg.db, keys=X, column="SYMBOL", keytype="ENSEMBL", multiVals="first") } )

ENSBTAG00000019313 ENSBTAG00000016711 ENSBTAG00000003877 ENSBTAG00000044714 ENSBTAG00000020940 ENSBTAG00000020939 ENSBTAG00000001656 ENSBTAG00000045788 ENSBTAG00000006536 
           "ZMIZ1"             "PPIF"          "ZCCHC24"                 NA           "ANXA11"            "PLAC9"          "TMEM254"                 NA             "CL46" 

When looking closer at these genes in Ensembl, I notice that they are all located "to the right" of the SNP location on the forward strand and that some of them are on the Forward strand (e.g. ZMIZ1 and PPIF), while others are on the Reverse strand (e.g. ZCCHC24 and PLAC9):

https://oct2018.archive.ensembl.org/Bos_taurus/Location/View?db=core;g=ENSBTAG00000013264;r=28:34948822-35454825;t=ENSBTAT00000046662

This seems a bit confusing. Shouldn't the variant be considered upstream relative to the two genes on the Forward strand, and downstream relative to the genes in the Reverse strand?

How should the PRECEDEID gene list be interpreted, more precisely?

Ngôn ngữ chính
R
Star
32
Fork
21
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của Bioconductor/VariantAnnotation

Tất cả issue của Bioconductor/VariantAnnotation

Issue tương tự

Thêm issue về R

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.