Experiment / E5PGRBQW9Sort-Seq / Flow-Seq MPRA

HEK293T 25-bp MAE-Seq fluorescence-selected enhancer library

Start from Scratch: Precisely Identify Massive Active Enhancers by Sequencing

A random 25-bp human genomic fragment library was cloned into a mini-promoter mCherry reporter plasmid and transfected into HEK293T cells. Fluorescence-selected cells were amplified and sequenced to identify active core enhancer loci; the released table is the 9,771-site positive enhancer catalog.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

MAE-Seq used a plasmid reporter based on pMX-GFP, modified to carry a mini-promoter and mCherry (pMX-mP-mCherry). Synthetic random DNA fragments of 20–60 bp were evaluated, with 25 bp selected for the final library. HEK293T cells at 70–80% confluence were transfected with 4 micrograms library DNA per 1 million cells using FuGENE HD; fluorescent cells were collected 24 hours later, vector-flanking PCR amplified, and sequenced on an Illumina HiSeq 2000. This is a fluorescence-selection/output-library assay rather than a barcode-level RNA/DNA ratio assay.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (23 of 23)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 23 definitions
element_id
Stable package-generated identifier for each retained N25 enhancer site.
source_peak_id
Original PeakID from the GEO/ChIPseek annotation; blank for one valid source row without a PeakID.
chromosome
hg19 chromosome or alternate/unplaced contig containing the site.
start_hg19_0based
0-based start coordinate from the source BED-derived annotation.
end_hg19_0based_exclusive
0-based exclusive end coordinate from the source annotation; each interval is 25 bp.
length_bp
Length of the mapped N25 interval in base pairs.
strand
Strand reported for the mapped fragment.
peak_score
Source Peak Score annotation; all retained rows have value 1 and it is not a quantitative MPRA activity effect.
annotation
Broad genomic annotation assigned by ChIPseek.
detailed_annotation
Detailed genomic or repeat annotation assigned by ChIPseek.
distance_to_tss_bp
Signed distance in base pairs to the nearest annotated transcription start site.
nearest_promoter_id
Nearest promoter transcript identifier reported by the source annotation.
entrez_id
Entrez Gene identifier for the nearest annotated gene, when available.
nearest_unigene
Nearest UniGene cluster identifier, when available.
nearest_refseq
Nearest RefSeq transcript identifier, when available.
nearest_ensembl
Nearest Ensembl gene identifier, when available.
gene_name
Nearest gene name reported by the source annotation.
gene_alias
Aliases for the nearest gene, when available.
gene_description
Description of the nearest gene, when available.
gene_type
Nearest gene biotype reported by the source annotation.
ncbi_link
Source NCBI nucleotide link for the nearest RefSeq transcript.
uniprot_link
Source UniProt query link for the nearest gene.
ucsc_genome_browser
Source UCSC hg19 browser link for the site and its surrounding region.

Quality control

The authors trimmed adapters and homologous arms with cutadapt v1.16, retained reads in which more than 50% of bases had Phred quality >=20 using fastx-toolkit, removed duplicates with FastUniq, selected 25-bp reads, aligned them to hg19 with Bowtie2 v2.1.0 (-D 20 -R 3 -N 0 -L 17 -i S,1,0.50 --end-to-end), and retained uniquely mapped reads without insertions or deletions. The paper reports 7,412,991 raw reads, 7,412,979 after trimming, 7,406,623 after quality filtering, 532,017 after duplicate removal, 22,529 unique mapped reads, 12,412 deduplicated unique mapped reads, and 9,771 unique sites. Package QC removed the trailing empty line from the GEO annotation export; all 9,771 remaining records have valid 25-bp intervals and unique coordinates. Valid hg19 alternate/unplaced contigs were retained.

Curation notes

The sole GEO sample is GSM3713164 (N25 fragments), representing the fluorescence-selected output library. The public supplement is an annotated positive-site catalog, not a quantitative MPRA count matrix: no barcode-level DNA/RNA counts, allelic contrasts, log2 activity scores, or per-replicate measurements are available. The source Peak Score is 1 for every row and should be treated only as an annotation field. One valid 25-bp interval lacks a source PeakID and is retained under a package-generated element_id. Ten records map to valid hg19 alternate or unplaced contigs and are retained rather than discarded. HEK293T was resolved to Cellosaurus CVCL:0063.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.