代码之家  ›  专栏  ›  技术社区  ›  claudiadast

如何将列更改为基于分隔符的值列表

r
  •  0
  • claudiadast  · 技术社区  · 7 年前

    我目前有一个数据框,看起来像这样:

    SampleID    Chrom   Start       End         ID
    HSB275      chr1    243216377   243219494   ENST00000366542|ENSG00000143702|protein_coding|protein_coding,chr1,243216377,243219494;ENST00000366543|ENSG00000143702|protein_coding|protein_coding,chr1,243216377,243219494
    HSB274      chr10   952208      979839      ENST00000381466|ENSG00000205740|antisense|processed_transcript,chr10,971146,979839
    HSB272      chr10   1046378     1047984     ENST00000381344|ENSG00000067064|protein_coding|protein_coding,chr10,1046378,1047984;ENST00000491735|ENSG00000067064|processed_transcript|protein_coding,chr10,1046378,1047984;ENST00000427898|ENSG00000067064|protein_coding|protein_coding,chr10,1046378,1047984
    HSB481      chr11   654157      655184      ENST00000527170|ENSG00000177030|nonsense_mediated_decay|protein_coding,chr11,654157,655184
    

    我想做的是降低成本 ID 列中的“ENSGXXXXXXX”值列表,如果每行有多个值,则这些值由“,”分隔,因此它看起来像 Genes

    预期结果:

    SampleID    Chrom   Start       End         Genes
    HSB275      chr1    243216377   243219494   ENSG00000143702,ENSG00000143702
    HSB274      chr10   952208      979839      ENSG00000205740
    HSB272      chr10   1046378     1047984     ENSG00000067064,ENSG00000067064,ENSG00000067064
    HSB481      chr11   654157      655184      ENSG00000177030
    
    5 回复  |  直到 7 年前
        1
  •  2
  •   Jaap    7 年前

    您没有固定的分隔符,但使用 strpslit 我们可以分两部分 ID , ; , | ),则对于每个元素,仅保留以“ENSG”开头的值,并删除其他值。

    sapply(strsplit(df$ID, ",|\\||;"), 
              function(x) toString(grep("^ENSG", x, value = TRUE)))
    
    
    #[1] "ENSG00000143702, ENSG00000143702"                 
    #[2] "ENSG00000205740"                                  
    #[3] "ENSG00000067064, ENSG00000067064, ENSG00000067064"
    #[4] "ENSG00000177030"      
    
        2
  •  0
  •   Maurits Evers    7 年前

    tidyverse

    library(tidyverse)
    df %>%
        mutate(Genes = map_chr(str_split(ID, ";"), ~toString(map(str_split(.x, "\\|"), 2)))) %>%
        select(-ID)
    #  SampleID Chrom     Start       End
    #1   HSB275  chr1 243216377 243219494
    #2   HSB274 chr10    952208    979839
    #3   HSB272 chr10   1046378   1047984
    #4   HSB481 chr11    654157    655184
    #                                              Genes
    #1                  ENSG00000143702, ENSG00000143702
    #2                                   ENSG00000205740
    #3 ENSG00000067064, ENSG00000067064, ENSG00000067064
    #4                                   ENSG00000177030
    

    样本数据

    df <- read.table(text =
        "SampleID    Chrom   Start       End         ID
    HSB275      chr1    243216377   243219494   ENST00000366542|ENSG00000143702|protein_coding|protein_coding,chr1,243216377,243219494;ENST00000366543|ENSG00000143702|protein_coding|protein_coding,chr1,243216377,243219494
    HSB274      chr10   952208      979839      ENST00000381466|ENSG00000205740|antisense|processed_transcript,chr10,971146,979839
    HSB272      chr10   1046378     1047984     ENST00000381344|ENSG00000067064|protein_coding|protein_coding,chr10,1046378,1047984;ENST00000491735|ENSG00000067064|processed_transcript|protein_coding,chr10,1046378,1047984;ENST00000427898|ENSG00000067064|protein_coding|protein_coding,chr10,1046378,1047984
    HSB481      chr11   654157      655184      ENST00000527170|ENSG00000177030|nonsense_mediated_decay|protein_coding,chr11,654157,655184", header = T)
    
        3
  •  0
  •   A. Suliman    7 年前
    library(dplyr)
    library(stringr) #str_extract_all
    df %>% group_by(SampleID) %>% #Use rowwise() if you do not like group_by
           mutate(Genes = paste(str_extract_all(ID, 'ENSG\\d+',simplify = T),collapse = ',')) %>% 
           select(-ID)
    
    # A tibble: 4 x 5
    # Groups:   SampleID [4]
    SampleID Chrom     Start       End Genes                                          
    <fct>    <fct>     <int>     <int> <chr>                                          
    1 HSB275   chr1  243216377 243219494 ENSG00000143702,ENSG00000143702                
    2 HSB274   chr10    952208    979839 ENSG00000205740                                
    3 HSB272   chr10   1046378   1047984 ENSG00000067064,ENSG00000067064,ENSG00000067064
    4 HSB481   chr11    654157    655184 ENSG00000177030  
    
        4
  •  0
  •   Nautica    7 年前

    我的尝试:

    genes %>% 
      mutate_at(vars(ID), funs(str_extract_all(., "ENSG[:digit:]*") %>% 
                                         str_replace_all("c|\"|\\(|\\)", "")))
    
        # A tibble: 4 x 5
      SampleID Chrom     Start       End ID                                               
      <chr>    <chr>     <dbl>     <dbl> <chr>                                            
    1 HSB275   chr1  243216377 243219494 ENSG00000143702, ENSG00000143702                 
    2 HSB274   chr10    952208    979839 ENSG00000205740                                  
    3 HSB272   chr10   1046378   1047984 ENSG00000067064, ENSG00000067064, ENSG00000067064
    4 HSB481   chr11    654157    655184 ENSG00000177030   
    

    这将查找与匹配的任何模式 ENSG<any length of numeric characters> ,然后将列表强制为相关字符串的向量,并整理所有不需要的字符。

    虽然我个人的精神是整洁的数据,但我会将每个“ID”放在一个单独的列中,复制相关的SampleID/Chrom/Start/End数据。

        5
  •  0
  •   Gwang-Jin Kim    7 年前

    数据

    df <- read.table(text =
        "SampleID    Chrom   Start       End         ID
    HSB275      chr1    243216377   243219494   ENST00000366542|ENSG00000143702|protein_coding|protein_coding,chr1,243216377,243219494;ENST00000366543|ENSG00000143702|protein_coding|protein_coding,chr1,243216377,243219494
    HSB274      chr10   952208      979839      ENST00000381466|ENSG00000205740|antisense|processed_transcript,chr10,971146,979839
    HSB272      chr10   1046378     1047984     ENST00000381344|ENSG00000067064|protein_coding|protein_coding,chr10,1046378,1047984;ENST00000491735|ENSG00000067064|processed_transcript|protein_coding,chr10,1046378,1047984;ENST00000427898|ENSG00000067064|protein_coding|protein_coding,chr10,1046378,1047984
    HSB481      chr11   654157      655184      ENST00000527170|ENSG00000177030|nonsense_mediated_decay|protein_coding,chr11,654157,655184", header = T)
    

    我的解决方案

    我定义了两个函数,这两个函数将来可能会有用,并且可以解决此任务:

    第一,, extract_matches ,提取中每个元素的所有匹配项 str.vec gregexpr 它只返回匹配的位置信息。

    第二,, extract_matches_aggregating sep= . 这取决于 提取匹配项 .

    您可以使用这两个函数提取所有ENSG ID,并通过“,”链接它们。

    extract_matches <- function(pattern, str.vec) {
      Map(function(m, s) substring(s, m, m + attr(m, "match.length") - 1), gregexpr(pattern, str.vec), str.vec)
    }
    
    extract_matches_aggregating <- function(pattern, str.vec, sep = "; ") {
      sapply(extract_matches(pattern, str.vec), function(res_vec) {
               paste(res_vec, collapse = sep)})
    }
    
    df$ID <- extract_matches_aggregating(pattern = "ENSG\\d+", str.vec = df$ID, sep = ", ")
    

    df 那么:

      SampleID Chrom     Start       End
    1   HSB275  chr1 243216377 243219494
    2   HSB274 chr10    952208    979839
    3   HSB272 chr10   1046378   1047984
    4   HSB481 chr11    654157    655184
                                                     ID
    1                  ENSG00000143702, ENSG00000143702
    2                                   ENSG00000205740
    3 ENSG00000067064, ENSG00000067064, ENSG00000067064
    4                                   ENSG00000177030
    

    在大型表上,此解决方案比使用 strsplit 和 sapply 和 lapply