Experiment / E1HG0ROWROther

BiT-STARR-seq allele-specific regulatory activity in HUVECs with caffeine versus water control

Characterization of caffeine response regulatory variants in vascular endothelial cells

An episomal BiT-STARR-seq library tested 43,556 SNP-centered 200-nt regulatory targets, with both alleles represented in forward and reverse construct orientations. HUVECs received 1.16 × 10^-3 M caffeine or water vehicle after transfection, with six biological replicates per condition; the table summarizes condition-specific ASE and caffeine-versus-water cASE results.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

1.16 × 10^-3 M caffeine for 24 h versus water vehicle control

The common library contained 43,556 SNP-centered target regions, 87,112 allele-specific constructs, and up to two construct orientations per allele; 1,676 computational negative-control sequences were included. DNA library was transfected into Lonza HUVECs with the Lonza Nucleofector X platform (DS-120, P5 solution), caffeine was added after transfection, cells were incubated 24 h, and RNA was sequenced on an Illumina NextSeq 500. Reads were UMI-filtered for the RDHBVDHBVD pattern, deduplicated, aligned to hg19, and converted to allele/direction counts. QuASAR-MPRA was used for condition-specific ASE and the laboratory's Delta-AST method for cASE.

BiT-STARR-seq (Biallelic Targeted STARR-seq): an episomal, oligo-synthesized self-transcribing reporter assay in which 200-nt SNP-centered inserts were synthesized for reference and alternate alleles and tested in forward and reverse orientations. Reporter RNA and input DNA were sequenced, and allele-specific effects were estimated from allele counts while accounting for DNA-library imbalance.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (45 of 45)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 45 definitions
identifier
Original study SNP/direction-pair identifier, formatted as rsID_fw or rsID_rv.
variant_id
The study rsID extracted from identifier.
direction
Reporter construct orientation label from the study: fw or rv.
chromosome
Primary hg19 chromosome for the rsID, when resolved.
position_hg19_0based
0-based hg19 genomic position, when resolved.
position_hg19_1based
1-based hg19 genomic position, when resolved.
coordinate_source
Source used for the added hg19 coordinate: the older Zenodo S4 coordinate bridge, Ensembl GRCh37 variation mapping, or unresolved.
geo_count_coordinate_match
Whether at least one GEO-deposited allele-count row occurs at the resolved chromosome and 1-based position.
grch37_allele_string
Ensembl GRCh37 allele string when available; it is not a re-called allele assignment for the MPRA.
geo_ref_alleles_observed
Distinct non-dot REF strings observed in the twelve GEO count files at the resolved position.
geo_alt_alleles_observed
Distinct non-dot ALT strings observed in the twelve GEO count files at the resolved position; may be NA when no alternate read was observed.
geo_count_rows_at_coordinate
Number of GEO count rows across all twelve libraries at the resolved position.
jaspar_motif_count
Number of distinct JASPAR PWMScan motif IDs overlapping the rsID in the study's motif annotation.
jaspar_motif_ids
Semicolon-separated JASPAR PWMScan motif IDs overlapping the rsID; NA means no motif annotation in supplementary file 2.
caffeine_effect_estimate
Published QuASAR-MPRA meta_estimate for the caffeine condition.
caffeine_effect_se
Published standard error of the caffeine ASE effect estimate.
caffeine_replicates_with_nonzero_counts
Number of the six caffeine replicates containing nonzero allele counts (published n.x).
caffeine_dna_reference_proportion
Published reference-allele proportion in the DNA library for caffeine analysis.
caffeine_z
Published caffeine ASE meta z-score.
caffeine_p_value
Published nominal caffeine ASE p-value.
caffeine_ase_fdr
Published Benjamini–Hochberg-adjusted caffeine ASE p-value (meta_padj.x).
published_new_padj
Published new_padj field from final supplementary file 4, retained without recomputation.
caffeine_ase_fdr10_significant
true when caffeine_ase_fdr < 0.10; otherwise false.
water_effect_estimate
Published QuASAR-MPRA meta_estimate for the water vehicle condition.
water_effect_se
Published standard error of the water ASE effect estimate.
water_replicates_with_nonzero_counts
Number of the six water replicates containing nonzero allele counts (published n.y).
water_dna_reference_proportion
Published reference-allele proportion in the DNA library for water analysis.
water_z
Published water ASE meta z-score.
water_p_value
Published nominal water ASE p-value.
water_ase_fdr
Published Benjamini–Hochberg-adjusted water ASE p-value (meta_padj.y).
water_ase_fdr10_significant
true when water_ase_fdr < 0.10; otherwise false.
conditional_ase_z
Published caffeine-versus-water conditional ASE Delta-Z statistic (case_z).
conditional_ase_p_value
Published nominal conditional ASE p-value (case_p).
conditional_ase_fdr
Published Benjamini–Hochberg-adjusted conditional ASE p-value (case_padj).
conditional_ase_fdr5_significant
true when conditional_ase_fdr < 0.05; otherwise false.
target_base_mean
Author-provided DESeq2 normalized mean activity for the coordinate-matched target from supplementary file 1.
target_activity_log2_fold_change
Author-provided DESeq2 caffeine-versus-control log2 fold change for the coordinate-matched target.
target_activity_lfc_se
Author-provided standard error of the target activity log2 fold change.
target_activity_stat
Author-provided DESeq2 test statistic for target activity.
target_activity_p_value
Author-provided nominal p-value for target activity.
target_activity_fdr
Author-provided Benjamini–Hochberg-adjusted p-value for target activity.
target_activity_fdr10_significant
true when the coordinate-matched target activity FDR is <0.10; NA means no unambiguous supplementary-file-1 coordinate join.
qc_pass_caffeine_4of6
true when the caffeine identifier has n.x >= 4; this is the paper's replicate-support threshold.
qc_pass_water_4of6
true when the water identifier has n.y >= 4; this is the paper's replicate-support threshold.
qc_pass_any_condition
Row-level inclusion flag; true when at least one condition passes the 4-of-6 replicate-support threshold.

Quality control

The authors demultiplexed with bcl2fastq, aligned to hg19 with HISAT2, removed reads with short or non-matching RDHBVDHBVD UMIs, deduplicated with UMItools, and generated allele counts with samtools mpileup followed by bcftools query. QuASAR-MPRA modeled allelic imbalance using the DNA-library allele proportion; six replicates were available per condition. The study required each identifier to have nonzero counts in at least 4 of 6 replicates before multiple-testing correction, used Benjamini–Hochberg FDR <10% for ASE and <5% for cASE, and tested 23,814 SNP/direction pairs with both conditions meeting the replicate threshold for cASE. table.csv excludes rows failing the 4-of-6 threshold in both conditions; rows passing only one condition are retained for that valid condition and have NA cASE fields.

Curation notes

This is one combined MPRA experiment because the paper used one shared BiT-STARR-seq library and analyzed six caffeine and six water HUVEC libraries as condition-specific ASE and cASE contrasts; the two library batches are represented by BST5 and BST11 sample prefixes. The final eLife supplementary file 4 is the authoritative source for all effect, p-value, and FDR columns in table.csv. It does not include genomic coordinates, so coordinates are added by rsID using the retained 2022 Zenodo preprint S4 bridge when available and Ensembl's GRCh37 variation mapping for remaining primary-assembly rsIDs; alternate-patch mappings and six unresolved rsIDs are left NA. GEO REF/ALT columns are explicitly labeled observations from deposited count rows and are not re-called variant alleles. Supplementary-file-1 target activity is joined by exact hg19 coordinate and is therefore an annotation rather than a direction-resolved reanalysis. The GEO family metadata and filenames identify the batch-2 libraries as BST11HC1/BST11HW1, whereas the author supplementary QC workbook labels the corresponding rows BST11HC4/BST11HW4; the GEO identifiers are retained as authoritative. The package retains the final eLife supplementary tables and the post-UMI-filter/deduplication GEO allele-count outputs, but not raw sequencing reads.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.