Targeted AR binding-site STARR-seq in LNCaP cells
Functional mapping of androgen receptor enhancer activityA capture-based STARR-seq library of clinical androgen-receptor binding regions was electroporated into human LNCaP prostate-cancer cells and measured after androgen stimulation or vehicle treatment. This package summarizes the public GEO class BED files for 3,230 LNCaP-overlapping regions together with three DHT, three EtOH, and one plasmid-library input BigWig track.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
10 nM dihydrotestosterone (DHT) for 4 h versus 100% EtOH vehicle for 4 h
The second-generation hSTARR-ORI plasmid contained 500–800 bp genomic fragments captured with custom Agilent probes spanning 700-bp target regions and cloned downstream of a minimal promoter. More than 1.3 × 10^8 LNCaP cells per biological replicate were electroporated with the library; after overnight recovery and charcoal-stripped-serum culture, cells were treated approximately 72 h later with 10 nM DHT or EtOH for 4 h. Poly(A)-selected STARR RNA was reverse-transcribed with a gene-specific primer and amplified by junction PCR; the plasmid library was separately amplified as the input. Paired-end 150-bp Illumina HiSeq 4000 reads were mapped to hg19, and GEO released RPKM-normalized extended-read BigWig tracks for GSM4565628–GSM4565630 (DHT), GSM4565631–GSM4565633 (EtOH), and GSM4565634 (library input). The processed table reports mean signal over each 700-bp BED interval, input-normalized log2 activity summaries using a +1 RPKM pseudocount, and a DHT-versus-EtOH log2 fold-change; these derived scores are not a reimplementation of the authors’ DESeq2 count model.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (22 of 22)
| Row | ||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ||||||||||||||||||||||
| 2 | ||||||||||||||||||||||
| 3 | ||||||||||||||||||||||
| 4 | ||||||||||||||||||||||
| 5 | ||||||||||||||||||||||
| 6 | ||||||||||||||||||||||
| 7 | ||||||||||||||||||||||
| 8 | ||||||||||||||||||||||
| 9 | ||||||||||||||||||||||
| 10 | ||||||||||||||||||||||
| 11 | ||||||||||||||||||||||
| 12 | ||||||||||||||||||||||
| 13 | ||||||||||||||||||||||
| 14 | ||||||||||||||||||||||
| 15 | ||||||||||||||||||||||
| 16 | ||||||||||||||||||||||
| 17 | ||||||||||||||||||||||
| 18 | ||||||||||||||||||||||
| 19 | ||||||||||||||||||||||
| 20 | ||||||||||||||||||||||
| 21 | ||||||||||||||||||||||
| 22 | ||||||||||||||||||||||
| 23 | ||||||||||||||||||||||
| 24 | ||||||||||||||||||||||
| 25 | ||||||||||||||||||||||
| 26 | ||||||||||||||||||||||
| 27 | ||||||||||||||||||||||
| 28 | ||||||||||||||||||||||
| 29 | ||||||||||||||||||||||
| 30 | ||||||||||||||||||||||
| 31 | ||||||||||||||||||||||
| 32 | ||||||||||||||||||||||
| 33 | ||||||||||||||||||||||
| 34 | ||||||||||||||||||||||
| 35 | ||||||||||||||||||||||
| 36 | ||||||||||||||||||||||
| 37 | ||||||||||||||||||||||
| 38 | ||||||||||||||||||||||
| 39 | ||||||||||||||||||||||
| 40 | ||||||||||||||||||||||
| 41 | ||||||||||||||||||||||
| 42 | ||||||||||||||||||||||
| 43 | ||||||||||||||||||||||
| 44 | ||||||||||||||||||||||
| 45 | ||||||||||||||||||||||
| 46 | ||||||||||||||||||||||
| 47 | ||||||||||||||||||||||
| 48 | ||||||||||||||||||||||
| 49 | ||||||||||||||||||||||
| 50 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 22 definitions
- element_id
- Unique ARBS identifier from the GEO BED files.
- chromosome
- Chromosome name for the tested region.
- start_hg19
- 0-based inclusive start coordinate on hg19 from the BED file.
- end_hg19
- 0-based exclusive end coordinate on hg19 from the BED file.
- published_enhancer_class
- Author/GEO class: induced, constitutively_active, or nonactive.
- element_length_bp
- Interval length in base pairs; all retained regions are 700 bp.
- library_input_mean_rpkm
- Mean RPKM-normalized extended-read signal across the interval in the plasmid-library input track GSM4565634.
- dht_rep1_mean_rpkm
- Mean RPKM signal across the interval for DHT biological replicate 1 (GSM4565628).
- dht_rep2_mean_rpkm
- Mean RPKM signal across the interval for DHT biological replicate 2 (GSM4565629).
- dht_rep3_mean_rpkm
- Mean RPKM signal across the interval for DHT biological replicate 3 (GSM4565630).
- dht_mean_rpkm
- Arithmetic mean of the three DHT replicate interval signals.
- dht_sd_rpkm
- Sample standard deviation of the three DHT replicate interval signals.
- dht_activity_log2_input
- Derived log2 input-normalized DHT activity, log2((DHT mean RPKM + 1)/(library input RPKM + 1)).
- etoh_rep1_mean_rpkm
- Mean RPKM signal across the interval for EtOH biological replicate 1 (GSM4565631).
- etoh_rep2_mean_rpkm
- Mean RPKM signal across the interval for EtOH biological replicate 2 (GSM4565632).
- etoh_rep3_mean_rpkm
- Mean RPKM signal across the interval for EtOH biological replicate 3 (GSM4565633).
- etoh_mean_rpkm
- Arithmetic mean of the three EtOH replicate interval signals.
- etoh_sd_rpkm
- Sample standard deviation of the three EtOH replicate interval signals.
- etoh_activity_log2_input
- Derived log2 input-normalized EtOH activity, log2((EtOH mean RPKM + 1)/(library input RPKM + 1)).
- dht_vs_etoh_log2fc
- Derived DHT-versus-EtOH log2 fold-change, log2((DHT mean RPKM + 1)/(EtOH mean RPKM + 1)); not the authors’ DESeq2 statistic.
- qc_pass
- Boolean indicating that the row passed package interval, coverage, finite-signal, and positive-input QC.
- source_accession
- GEO series accession supplying the tracks and published enhancer-class intervals.
Quality control
The authors mapped STARR-seq reads with BWA-MEM 0.7.17, removed reads with indels or MAPQ <60, discarded ENCODE blacklist regions (ENCFF001TDO), generated RPKM-normalized extended-read tracks with bamCoverage, and used DESeq2 for differential enhancer activity. Their published calls defined induced enhancers as DHT/EtOH log2 fold-change >1 with adjusted p-value <0.05, constitutive enhancers as plasmid-normalized log2 activity >1 in both conditions without DHT induction, and inactive regions as low plasmid-normalized activity with minimal induction; replicate Pearson correlations were reported as 0.84–0.99. Package QC required the published BED interval to be exactly 700 bp, complete coverage in all seven supplied tracks, finite signal values, and positive plasmid-library input. Five zero-input regions were excluded, leaving 3,225 rows; valid nonactive and non-significant regions were retained.
Curation notes
The public GSE151064 processed files expose 465 constitutively active, 286 induced, and 2,479 nonactive LNCaP-overlapping ARBS intervals. The paper discusses a larger capture design that also included clinical-specific ARBS, non-AR positive controls, and ARE-motif-only controls, but the complete coordinate table for those additional targets is not separately released in this GEO series; the table therefore stays restricted to the released BED intervals. Five regions with zero GSM4565634 library-input signal were omitted from table.csv but remain in the raw BED files. The library input is a single shared track, so the input-normalized RPKM scores are compact derived summaries rather than original count-level DESeq2 results. GEO labels the line simply LNCaP; ATCC CRL-1740 corresponds to LNCaP clone FGC, resolved here to Cellosaurus CVCL:1379. The paper also performed luciferase and CRISPRi validation, including a separate somatic-SNV luciferase assay; this experiment package contains only the targeted STARR-seq data.