Variant and regulatory-fragment lentiMPRA screen in C4-2B cells
Integrative identification of non-coding regulatory regions driving metastatic prostate cancerA lentiMPRA library of 3,665 candidate regulatory sequence elements, scrambled controls, reference sequences, and patient-observed mutant variants was assayed in human C4-2B metastatic prostate-cancer cells. Three biological replicates were profiled through gDNA input and RNA output barcode counts, and logistic regression produced reference-versus-scrambled and mutant-versus-reference enrichment statistics.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
Basal / Untreated
The CRS library was cloned into the pLS-SceI lentiviral reporter backbone with a minimal promoter and 15-bp barcodes. C4-2B cells were infected at a target of 100 integrations per barcode in three biological replicates; DNA and RNA were co-extracted, RNA was DNase-treated and reverse-transcribed, and UMI-tagged DNA/RNA libraries were sequenced. GEO describes PEAR merging, Bowtie2 alignment, unique barcode-to-fragment assignment, UMI-based counting, and logistic regression of reference, mutant, and scrambled sequences.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (28 of 28)
| Row | ||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ||||||||||||||||||||||||||||
| 2 | ||||||||||||||||||||||||||||
| 3 | ||||||||||||||||||||||||||||
| 4 | ||||||||||||||||||||||||||||
| 5 | ||||||||||||||||||||||||||||
| 6 | ||||||||||||||||||||||||||||
| 7 | ||||||||||||||||||||||||||||
| 8 | ||||||||||||||||||||||||||||
| 9 | ||||||||||||||||||||||||||||
| 10 | ||||||||||||||||||||||||||||
| 11 | ||||||||||||||||||||||||||||
| 12 | ||||||||||||||||||||||||||||
| 13 | ||||||||||||||||||||||||||||
| 14 | ||||||||||||||||||||||||||||
| 15 | ||||||||||||||||||||||||||||
| 16 | ||||||||||||||||||||||||||||
| 17 | ||||||||||||||||||||||||||||
| 18 | ||||||||||||||||||||||||||||
| 19 | ||||||||||||||||||||||||||||
| 20 | ||||||||||||||||||||||||||||
| 21 | ||||||||||||||||||||||||||||
| 22 | ||||||||||||||||||||||||||||
| 23 | ||||||||||||||||||||||||||||
| 24 | ||||||||||||||||||||||||||||
| 25 | ||||||||||||||||||||||||||||
| 26 | ||||||||||||||||||||||||||||
| 27 | ||||||||||||||||||||||||||||
| 28 | ||||||||||||||||||||||||||||
| 29 | ||||||||||||||||||||||||||||
| 30 | ||||||||||||||||||||||||||||
| 31 | ||||||||||||||||||||||||||||
| 32 | ||||||||||||||||||||||||||||
| 33 | ||||||||||||||||||||||||||||
| 34 | ||||||||||||||||||||||||||||
| 35 | ||||||||||||||||||||||||||||
| 36 | ||||||||||||||||||||||||||||
| 37 | ||||||||||||||||||||||||||||
| 38 | ||||||||||||||||||||||||||||
| 39 | ||||||||||||||||||||||||||||
| 40 | ||||||||||||||||||||||||||||
| 41 | ||||||||||||||||||||||||||||
| 42 | ||||||||||||||||||||||||||||
| 43 | ||||||||||||||||||||||||||||
| 44 | ||||||||||||||||||||||||||||
| 45 | ||||||||||||||||||||||||||||
| 46 | ||||||||||||||||||||||||||||
| 47 | ||||||||||||||||||||||||||||
| 48 | ||||||||||||||||||||||||||||
| 49 | ||||||||||||||||||||||||||||
| 50 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 28 definitions
- sequence_id
- Exact sequence identifier from the GEO mutant-versus-WT result table; encodes the region and tested allele change.
- region_id
- Exact region identifier from the GEO result table, such as 988_chr1.
- variant_id
- Exact coordinate-and-allele label from the GEO result table, such as chr1_151526020_A_C; the source did not provide dbSNP rs identifiers.
- region_coordinate_hg38
- Broad GRCh38 interval of the tested regulatory fragment, reconstructed from the source sequence identifier.
- chromosome
- GRCh38 chromosome label, including the chr prefix.
- fragment_start_hg38
- Start coordinate of the tested regulatory fragment as encoded by GEO.
- fragment_end_hg38
- End coordinate of the tested regulatory fragment as encoded by GEO.
- variant_position_hg38
- Variant start position as encoded by GEO; for indels, this is the leftmost variant coordinate.
- reference_allele
- Reference allele string from the GEO sequence identifier.
- alternate_allele
- Alternate/mutant allele string from the GEO sequence identifier.
- variant_type
- SNV when both allele strings are one base; otherwise indel.
- region_length_bp
- Inclusive interval length calculated as fragment_end_hg38 minus fragment_start_hg38 plus one.
- mutant_vs_wt_logit_estimate
- Published logistic-regression enrichment estimate for the mutant allele relative to the reference allele; positive values indicate higher modeled mutant enrichment.
- mutant_vs_wt_logit_z
- Published z statistic for the mutant-versus-reference logistic-regression test.
- mutant_vs_wt_pvalue
- Published unadjusted p-value for the mutant-versus-reference test.
- mutant_vs_wt_qvalue
- Published multiple-testing-adjusted q-value for the mutant-versus-reference test.
- mutant_vs_wt_fdr_0_01
- TRUE when the published mutant-versus-reference q-value is below 0.01; derived package flag.
- wt_vs_scrambled_logit_estimate
- Published logistic-regression enrichment estimate for the exact sequence relative to its scrambled control, when that sequence is present in the WT-versus-scrambled GEO table.
- wt_vs_scrambled_logit_z
- Published z statistic for the exact sequence-versus-scrambled test, when available.
- wt_vs_scrambled_pvalue
- Published unadjusted p-value for the exact sequence-versus-scrambled test, when available.
- wt_vs_scrambled_qvalue
- Published multiple-testing-adjusted q-value for the exact sequence-versus-scrambled test, when available.
- wt_vs_scrambled_fdr_0_01
- TRUE when the exact sequence-versus-scrambled q-value is below 0.01; blank when no exact sequence match is available.
- wt_vs_scrambled_sequence_match
- TRUE when the exact sequence_id occurs in the deposited WT-versus-scrambled result table.
- is_cdrr
- TRUE when the fragment interval exactly matches a Candidate Driver Regulatory Region in article Supplementary Table S3.
- cdrr_consensus_function
- Consensus functional annotation for the matched CDRR from Supplementary Table S3, such as enhancer, promoter/enhancer, or UTR annotation.
- flanking_genes
- Comma-separated flanking gene symbols reported for the matched CDRR in Supplementary Table S3.
- cdrr_mutations_in_cohort
- Number of mutations in the mCRPC cohort for the matched CDRR, from Supplementary Table S3.
- cdrr_predicted_functional_mutations
- Number of predicted functional mutations for the matched CDRR, from Supplementary Table S3.
Quality control
The authors performed the lentiMPRA in biological triplicate and retained barcodes assigned to 358 fragments of interest and their scrambled controls with sufficient read counts (>25 reads per barcode) for downstream analysis. GEO processing used unique barcode assignment and UMI-based DNA/RNA counting before logistic regression. For this package, the mutant-versus-WT source table was additionally checked for non-empty identifiers and finite estimate, z, p, and q statistics; all 1,025 rows passed and were retained. The WT-versus-scrambled source table contains 358 complete rows, with 324 exact sequence identifiers shared with the mutant table. Rows were not filtered by statistical significance, so non-significant variant effects remain represented.
Curation notes
This child experiment represents the deposited C4-2B lentiMPRA only; the study's separate CRISPRi and CLIP-seq series were not packaged as MPRA experiments. The processed table is intentionally variant-centric and has one row per published mutant-versus-WT comparison (1,025 rows). Reference-versus-scrambled statistics are joined only when the exact GEO sequence identifier is shared (324 rows); blank joined values mean no exact match and are not zero. Region annotations are joined by exact GRCh38 interval to article Supplementary Table S3 (369 variant rows spanning 89 CDRRs). The source labels contain coordinates and allele strings but no rsIDs, so no external variant identifiers were inferred. Raw FASTQ/SRA reads were omitted; the GEO result tables, GEO family metadata, and small article supplements are retained in raw_data.