Experiment / E2WKYMR76Standard STARR-seq

Modified STARR-seq allelic enhancer screen in SH-SY5Y cells

Analysis of biased allelic enhancer activity of schizophrenia-linked common variants

The shared synthetic oligonucleotide library of schizophrenia-linked candidate variants was transfected into human SH-SY5Y neuroblastoma cells. Reference and alternative 200-bp allele fragments were tested in three SNP-centered contexts, with plasmid DNA input and RNA output quantified across three replicates.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Episomal modified STARR-seq using the hSTARR-seq_ORI vector (Addgene 99296). Each selected variant was represented by reference and alternative 200-bp oligonucleotides with the variant at -50, 0, or +50 bp from the sequence center; 15-bp cloning adapters were added. Plasmid-library DNA input and polyadenylated RNA output were sequenced as paired-end 150-bp libraries with three DNA and three SH-SY5Y RNA replicates. The published analysis used DADA2, limma-voom, and mpralm; this package retains allele/context-level counts and published eSNP and baaSNP statistics.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (48 of 48)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 48 definitions
variant_id
dbSNP rs identifier for the tested candidate variant.
chromosome
GRCh38 chromosome without a chr prefix.
position_hg38
1-based GRCh38 variant position.
allele1_base
Nucleotide label for reported allele 1 when resolved from the supplementary allele tables.
allele2_base
Nucleotide label for reported allele 2 when resolved from the supplementary allele tables.
allele_label_source
Source used for the nucleotide labels; blank means the source tables gave conflicting orientations.
fragment_id
Original Base_SNP fragment identifier from the GEO processed count table.
fragment_context
Synthetic oligonucleotide context, expressed as upstream and downstream bases around the SNP.
upstream_bp
Number of bases upstream of the SNP in the 200-bp fragment.
downstream_bp
Number of bases downstream of the SNP in the 200-bp fragment.
dna_allele1_rep1_count
Raw plasmid DNA input count for allele 1, replicate 1.
dna_allele1_rep2_count
Raw plasmid DNA input count for allele 1, replicate 2.
dna_allele1_rep3_count
Raw plasmid DNA input count for allele 1, replicate 3.
dna_allele2_rep1_count
Raw plasmid DNA input count for allele 2, replicate 1.
dna_allele2_rep2_count
Raw plasmid DNA input count for allele 2, replicate 2.
dna_allele2_rep3_count
Raw plasmid DNA input count for allele 2, replicate 3.
rna_allele1_rep1_count
Raw RNA output count for allele 1, replicate 1.
rna_allele1_rep2_count
Raw RNA output count for allele 1, replicate 2.
rna_allele1_rep3_count
Raw RNA output count for allele 1, replicate 3.
rna_allele2_rep1_count
Raw RNA output count for allele 2, replicate 1.
rna_allele2_rep2_count
Raw RNA output count for allele 2, replicate 2.
rna_allele2_rep3_count
Raw RNA output count for allele 2, replicate 3.
dna_allele1_mean_cpm
Mean allele-1 DNA input counts per million across the three DNA replicates, normalized within each replicate before averaging.
dna_allele2_mean_cpm
Mean allele-2 DNA input counts per million across the three DNA replicates, normalized within each replicate before averaging.
rna_allele1_mean_cpm
Mean allele-1 RNA output counts per million across the three RNA replicates, normalized within each replicate before averaging.
rna_allele2_mean_cpm
Mean allele-2 RNA output counts per million across the three RNA replicates, normalized within each replicate before averaging.
allele1_log2_rna_dna
Derived log2 RNA/DNA activity for allele 1 from mean CPM values with a 0.5-CPM pseudocount.
allele2_log2_rna_dna
Derived log2 RNA/DNA activity for allele 2 from mean CPM values with a 0.5-CPM pseudocount.
derived_allele2_minus_allele1_log2_activity
Derived allele-2 minus allele-1 difference of the CPM-based log2 RNA/DNA activities; not the paper's fitted mpralm statistic.
published_allele1_log2fc
Published Supplementary Data 2 limma-voom log2 fold-change for allele 1 output RNA versus input DNA.
published_allele1_pvalue
Published Supplementary Data 2 p-value for the allele-1 output-versus-input test.
published_allele1_fdr
Published Supplementary Data 2 Benjamini-Hochberg adjusted p-value for allele 1.
published_allele1_type
Published allele-1 activity class: eSNP, silencer, or inactive.
published_allele2_log2fc
Published Supplementary Data 2 limma-voom log2 fold-change for allele 2 output RNA versus input DNA.
published_allele2_pvalue
Published Supplementary Data 2 p-value for the allele-2 output-versus-input test.
published_allele2_fdr
Published Supplementary Data 2 Benjamini-Hochberg adjusted p-value for allele 2.
published_allele2_type
Published allele-2 activity class: eSNP, silencer, or inactive.
published_allelic_log2fc
Published Supplementary Data 3 mpralm log2 fold-change for the allelic comparison, signed according to the supplied A1/A2 labels.
published_allelic_pvalue
Published Supplementary Data 3 mpralm p-value for the allelic comparison.
published_allelic_t
Published Supplementary Data 3 mpralm t-statistic for the allelic comparison.
published_baa_label
Published Supplementary Data 3 baaSNP call: YES or NO.
eqtl_target_genes
Comma-separated candidate target genes from the study's integrated eQTL analysis; blank when not reported for the fragment.
hic_target_genes
Comma-separated candidate target genes from the study's integrated chromatin-interaction analysis; blank when not reported for the fragment.
allele1_qc_replicates_above_0_1_tpm
Number of available allele-1 DNA/RNA replicate columns exceeding the 0.1-TPM count threshold used for package QC.
allele2_qc_replicates_above_0_1_tpm
Number of available allele-2 DNA/RNA replicate columns exceeding the 0.1-TPM count threshold used for package QC.
dna_replicate_count
Number of DNA input replicates available in the GEO table.
rna_replicate_count
Number of RNA output replicates available in the GEO table.
qc_pass
TRUE for a row retained after the paired-allele package QC filter.

Quality control

The authors report read quality filtering and trimming, DADA2 ASV inference, paired-end merging, chimera removal, low-expression filtering, residual-log-expression normalization, hierarchical clustering, and D-statistic outlier assessment. Their enhancer calls used limma-voom with FDR <0.05 and |log2FC| >0.585; allelic effects used mpralm. For this package, the GEO-supplied filtered count table was additionally filtered at the fragment level: for each allele, counts had to exceed 0.1 TPM in >20% of the six available DNA/RNA replicate columns and in at least one DNA and one RNA replicate; both alleles had to pass. The 598 retained rows in table.csv pass this paired-allele QC; 363 of 961 GEO rows were excluded. The study reported no positive or negative control sequences.

Curation notes

The GEO supplementary table is a processed allele/context count table rather than raw FASTQ and is preserved in raw_data. The source count columns use allele1/allele2 names but do not contain nucleotide symbols; allele bases are taken from Supplementary Data 3 when present and otherwise from Supplementary Data 1, with allele_label_source exposing provenance. Supplementary Data 1 contains occasional reverse A1/A2 reports. Two GEO-only indel IDs (rs112345465 and rs34196118) occur in the count source but not Supplementary Data 1-3; they are retained only if paired QC passes, with coordinates and allele strings resolved from Ensembl and published result fields left blank. No positive/negative controls were included in the library.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.