A custom 11,999-construct library of 130–220 bp coding-exon fragments (candidate exonic enhancers, controls, and gnomAD-derived single- and multi-SNP variants) was cloned downstream of a constitutive promoter in a STARR-seq reporter and transfected into human K-562 cells. Five biological replicates were harvested 24 h later; reporter cDNA was quantified against input plasmid DNA.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Organism
Human
Taxonomy ID
NCBITaxon:9606
Biosample
CVCL:0004
Reference genome
GRCh38
Design focus
Region-focused
Region of interest
Not reported / not applicable
Perturbation & assay details
Basal / untreated K-562 culture; 24 h after electroporation
Exon-bounded 130–220 bp inserts were placed downstream of a constitutively active promoter, upstream of the reporter ORF and poly(A) signal. The library was synthesized by Twist Bioscience, assembled by Gibson cloning, and expanded at approximately 516× insert coverage. A total of 250 µg library plasmid was electroporated into 50 million K-562 cells through ten Neon transfections (1450 V, 10 ms, 3 pulses); two transfections were combined per culture to produce five biological replicates, with one empty-vector control replicate. Polyadenylated RNA was converted to cDNA and sequenced alongside input plasmid DNA on an Illumina NextSeq 2000. The released Zenodo summary includes five cDNA output count/CPM tracks (AVO5–AVO9).
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (49 of 49)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
Page 1 · 50 rows · More results available
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 49 definitions
element_id
Unique identifier for the synthesized library construct.
category
Original construct category from the authors' Category field.
category_group
Grouped construct category from the authors' Category2 field.
construct_class
Normalized interpretation of the library construct category.
parent_element_id
Shared three-part exon identifier used as the parent for variant constructs; blank for nonvariants.
predicted_tfbs_effect
Predicted transcription-factor-binding effect encoded by the source ID: gain or loss; blank for nonvariants.
variant_position_hg38
Single-SNP genomic position from the construct ID on GRCh38; blank for multi-SNP constructs and nonvariants.
ref_allele
Reference allele from a single-SNP construct ID; blank otherwise.
alt_allele
Alternate allele from a single-SNP construct ID; blank otherwise.
reference_element_id
Authors' mapped reference cEE construct for a single- or multi-SNP variant.
paired_multi_variant_id
Authors' mapped multi-SNP construct paired with a single-SNP construct.
chromosome
GRCh38 chromosome for the tested genomic fragment; blank for synthetic promoter/randomized controls.
start_hg38
GRCh38 zero-based inclusive start coordinate of the tested genomic fragment; blank for synthetic controls.
end_hg38
GRCh38 zero-based half-open end coordinate of the tested genomic fragment; blank for synthetic controls.
strand
Reference-strand orientation of the tested fragment.
sequence
Synthesized exon/control/variant DNA sequence used in the reporter library.
sequence_length
Length of sequence in nucleotides.
input_count_rep1
Raw input-DNA read count for AVO1ExonhDNA1.
input_cpm_rep1
CPM-normalized input-DNA count for AVO1ExonhDNA1.
input_count_rep2
Raw input-DNA read count for AVO2ExonhDNA2.
input_cpm_rep2
CPM-normalized input-DNA count for AVO2ExonhDNA2.
input_count_rep3
Raw input-DNA read count for AVO3ExonhDNA3.
input_cpm_rep3
CPM-normalized input-DNA count for AVO3ExonhDNA3.
plasmid_input_count
Raw maxiprep plasmid input count for AVO4Exonhmaxi4; used for the >=500 QC threshold.
plasmid_input_cpm
CPM-normalized maxiprep plasmid input count for AVO4Exonhmaxi4.
output_count_rep1
Raw reporter cDNA count for AVO5ExonhcDNA1.
output_cpm_rep1
CPM-normalized reporter cDNA count for AVO5ExonhcDNA1.
output_count_rep2
Raw reporter cDNA count for AVO6ExonhcDNA2.
output_cpm_rep2
CPM-normalized reporter cDNA count for AVO6ExonhcDNA2.
output_count_rep3
Raw reporter cDNA count for AVO7ExonhcDNA3.
output_cpm_rep3
CPM-normalized reporter cDNA count for AVO7ExonhcDNA3.
output_count_rep4
Raw reporter cDNA count for AVO8ExonhcDNA4.
output_cpm_rep4
CPM-normalized reporter cDNA count for AVO8ExonhcDNA4.
output_count_rep5
Raw reporter cDNA count for AVO9ExonhcDNA5.
output_cpm_rep5
CPM-normalized reporter cDNA count for AVO9ExonhcDNA5.
input_cpm_sd
Input_STD: source standard deviation across the three isolated input-DNA CPM measurements.
output_cpm_mean_reported_reps
Source Output_AVG activity metric, based on the AVO7–AVO9 output CPM tracks used by the authors' Log2FC calculation.
output_cpm_mean_all5
Mean CPM across all five released reporter cDNA output tracks, AVO5–AVO9; calculated here.
Number of total transcription-factor ChIP peaks annotated for the element.
num_k562_chip_peaks
Number of K-562 transcription-factor ChIP peaks annotated for the element.
num_total_starr_peaks
Number of total STARR-seq atlas peaks annotated for the element.
num_k562_starr_peaks
Number of K-562 STARR-seq atlas peaks annotated for the element.
dnase_overlap
Source DNase overlap indicator.
dnase_k562_score
Source K-562 DNase-seq score.
snp_vs_reference_padj
Benjamini–Hochberg adjusted Wald-test p-value for the SNP construct versus its mapped reference construct.
variant_delta_log2fc_vs_reference
Calculated log2fc_reported of the variant minus its mapped reference; populated when both constructs pass QC.
single_vs_multi_padj
Benjamini–Hochberg adjusted Wald-test p-value for the single-SNP versus paired multi-SNP comparison.
single_vs_multi_delta_log2fc
Calculated log2fc_reported of a single-SNP construct minus its paired multi-SNP construct; populated when both pass QC.
Quality control
Reads were aligned to the custom library with Bowtie2 using the --very-sensitive preset; unmapped and reverse-strand reads were removed with SAMtools -F 20; counts were normalized to CPM and cDNA/input log2 ratios were calculated. Following the released processing workflow, table entries required a numeric Log2FC, maxiprep plasmid input count >=500, and Input_STD <=20. Of 11,999 library constructs, 10,058 passed these filters. The article prose describes discarding Input STD >=20 whereas the released notebook uses <=20; no released row has exactly 20, so the retained set is unchanged.
Curation notes
The experiment combines a region-focused exon-enhancer/control screen with a large variant arm; variant annotations are retained rather than splitting the study into a second experiment. Category is the original label, while Category2 is the authors' grouped label (including EEK, EEKA, EEKAG, and EEKG under EEK). Sequences and coordinates were joined from the authors' STARR_AllSequences table and are represented as GRCh38 zero-based half-open intervals. G/L tokens in variant construct IDs were interpreted as the source's predicted TFBS gain/loss labels. Variant/reference and single-/multi-SNP fields come from the released Zenodo mappings and Wald-test tables. The complete Zenodo summary contains five cDNA output tracks (AVO5–AVO9), whereas the GEO processed summary exposes AVO7–AVO9; all five are retained here and the source Output_AVG/Log2FC fields are explicitly labeled. Raw sequence reads were not included.