A pooled allele-specific STARR-seq library containing 96 amplicons covering 101 prostate-cancer risk SNPs (202 allele sequences) was assayed in human LNCaP cells under the ethanol vehicle condition. This experiment represents the 0.1% ethanol control with two technical replicates and matched plasmid input controls; the processed table retains 100 SNP rows after condition-specific coverage/BAE QC.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Organism
Human
Taxonomy ID
NCBITaxon:9606
Biosample
CVCL:0395
Reference genome
hg19
Design focus
Variant-focused
Region of interest
Not reported / not applicable
Perturbation & assay details
0.1% ethanol vehicle control (ETH)
The authors used an episomal human STARR-seq vector and cloned pooled PCR amplicons containing the selected SNPs by In-Fusion recombination; the amplicons were generated from mixed DNA from 111 prostate cancer patients and placed downstream of a super core promoter. LNCaP cells received the plasmid library by Lipofectamine 3000, and polyA+ reporter RNA was isolated 24 h after transfection alongside plasmid DNA input controls. Reporter-specific cDNA and input DNA libraries were sonicated to 100–150 bp and sequenced as 100-bp paired-end Illumina libraries. Supplementary Data 4 reports two ETH vehicle technical replicates, two DHT replicates, and two plasmid-input replicates; the source BAE score is the log2 allele-ratio effect after test/input normalization.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (33 of 33)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
Page 1 · 50 rows · More results available
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 33 definitions
rsid
dbSNP identifier of the tested SNP.
allele_1
Allele symbol from the first allele row in Supplementary Data 4.
allele_2
Allele symbol from the second allele row in Supplementary Data 4.
eth_rep1_allele_1_read_count
Source-reported STARR-seq ETH read count for allele_1 in technical replicate 1.
eth_rep2_allele_1_read_count
Source-reported STARR-seq ETH read count for allele_1 in technical replicate 2.
eth_rep1_allele_2_read_count
Source-reported STARR-seq ETH read count for allele_2 in technical replicate 1.
eth_rep2_allele_2_read_count
Source-reported STARR-seq ETH read count for allele_2 in technical replicate 2.
input_rep1_allele_1_read_count
Source-reported plasmid-input read count for allele_1 in input replicate 1.
input_rep2_allele_1_read_count
Source-reported plasmid-input read count for allele_1 in input replicate 2.
input_rep1_allele_2_read_count
Source-reported plasmid-input read count for allele_2 in input replicate 1.
input_rep2_allele_2_read_count
Source-reported plasmid-input read count for allele_2 in input replicate 2.
bae_score
Source-reported ETH biased allelic enhancer (BAE) score, a log2 allele-ratio effect after normalizing the ETH test ratio to the plasmid-input ratio.
abs_bae_score
Absolute value of bae_score.
bae_effect_direction
Derived direction of the BAE score: allele_1_higher for positive, allele_2_higher for negative, or no_difference for zero.
bae_passes_abs_0_58
Derived boolean indicating whether absolute BAE is at least 0.58, the paper’s 1.5-fold cutoff.
enhancer_activity_based_selection
Source Supplementary Data 4 selection flag for the initial test/input enhancer-activity screen.
bae_score_based_selection
Source Supplementary Data 4 flag for the BAE-score selection step.
final_20_snps
Source Supplementary Data 4 flag identifying the final 20 SNPs selected by the study.
final_20_eqtl_gene
eQTL gene reported in the source row for a final-20 SNP; NA when not reported.
eqtl_gene
Candidate eQTL gene annotation from Supplementary Data 1.
chr_position_hg19
SNP coordinate from Supplementary Data 1, normalized from chr.position to chr:position on hg19.
position_in_genome
Source annotation of the SNP’s genomic context, such as intron or intergenic.
r2_with_lead_snp
Source linkage-disequilibrium R² with the listed lead SNP.
lead_snps
Lead prostate-cancer risk SNP or SNPs associated with the candidate locus.
lead_snp_position_hg19
hg19 position of the listed lead SNP.
eqtl_p_value
Source eQTL P value or semicolon-separated values for the candidate gene(s).
eqtl_p_value_classification
Source classification of the eQTL P value.
eqtl_fdr
Source eQTL false-discovery rate or semicolon-separated values.
kb_snp_to_tss
Source distance from the SNP to the target transcription start site in kilobases; semicolon-separated for multiple genes.
chip_seq_score
Source count of overlapping prostate-specific ChIP-seq/epigenomic signals.
amplicon_length_bp
Length in base pairs of the PCR amplicon used for the STARR-seq insert library.
snp_position_in_amplicon
Source SNP position within the amplicon in base pairs.
source_table
Supplementary Data 4 STARR-seq result table used for the reporter measurements.
Quality control
The paper mapped reads to 96 amplicons covering all 101 selected SNPs; mapped reads averaged 76% of raw reads (approximately 70% for polyA+ mRNA and 81% for plasmid DNA input), and each technical replicate pair had R² ≥ 0.95. The paper reports a median depth of 20,961 reads per allele (range 0–622,402), selected 56 SNPs with at least one allele showing a 1.5-fold test/input difference, and defined allele-dependent enhancer activity by |BAE| ≥ 0.58; 20 SNPs were in the final intersection. Package QC retained 100 of 101 SNP rows for this ETH condition by requiring a finite source ETH BAE score and nonzero aggregate ETH and input counts for both alleles; rs113812061 was excluded. Source selection flags are retained as annotations, and nonsignificant but QC-passing measured SNPs remain.
Curation notes
LNCaP was resolved to Cellosaurus CVCL:0395. The processed table is one row per SNP with the source allele order preserved as allele_1 and allele_2; Supplementary Data 4 does not provide a universal reference/alternate label for the two rows, so no ref/alt designation was inferred. The source BAE score is recorded once per SNP in the first allele row and is not recomputed from rounded display counts. The DHT and ETH conditions share the same plasmid-input controls and are represented as separate child experiments. Blank source fields are represented as NA in the processed CSV.