readVcf is Slow if ScanVcfParam which Regions is Lengthy
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 43/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- r
- Lĩnh vực
- data, performance
Hướng nghiên cứu
Bắt đầu bằng cách tái hiện chênh lệch thời gian giữa readVcf có và không có ScanVcfParam(which = goldStandards), sử dụng ví dụ GRanges bắt nguồn từ BED đã được báo cáo. Đọc các entry point của readVcf và ScanVcfParam, rồi so sánh hành vi với subsetByOverlaps. Công việc được xem là hoàn tất khi có một thay đổi đã được thống nhất trong mã hoặc tài liệu, giúp việc đọc VCF theo vùng trở nên thiết thực hoặc ghi lại rõ ràng workaround.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Recently, U.S. Food and Drug Administration, National Institute of Standards and Technology and Illumina researchers defined highly reproducible regions (H.R.R.s) of the human genome and made available BED files to define a set of regions in which variant callers consistently call variants on technical replicate samples, effectively defining a while-list for whole genome sequencing data. readVcf takes a long time if which of ScanVcfParam is specified. Importing the V.C.F. takes about five minutes if which not specified but I terminated it after one hour when which was specified.
library(rtracklayer)
goldStandards <- list.files("HRR/", "bed", full.names = TRUE)
goldStandards <- lapply(goldStandards, import.bed)
goldStandards <- unlist(goldStandards)
goldStandards <- reduce(goldStandards)
> goldStandards
GRanges object with 2988875 ranges and 0 metadata columns:
... ...
variants <- readVcf("DRAGENgermline.vcf.gz", param = ScanVcfParam(which = goldStandards)) # Stopped after one hour.
Importing the whole V.C.F. file into the session and then using subsetByOverlaps seems a reasonable and fast workaround. To best help other users, would a documentation change or code change be better to make this user experience nicer?
- Ngôn ngữ chính
- R
- Star
- 32
- Fork
- 21
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của Bioconductor/VariantAnnotation
-
refactor the package Đang mở
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
Bioconductor/VariantAnnotation#115 · 4 bình luận ·
-
Bioconductor/VariantAnnotation#114 · 1 reaction · 2 người được giao ·
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Bioconductor/VariantAnnotation#113 · 2 bình luận ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 35/100
-
ensemblVEP::parseCSQToGRanges Đang mở
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 30/100
Bioconductor/VariantAnnotation#88 · 2 bình luận ·
Tất cả issue của Bioconductor/VariantAnnotation
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
briandconnelly/airnow#9 ·
-
Copy cohorts to keep old cohorts Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
OHDSI/CohortConstructor#774 ·
-
pre-review R TeX Track: 5 (DSAIS)
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 60/100
openjournals/joss-reviews#11330 · 7 bình luận ·
-
Release autosync 0.1.1 Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100