Experiment / E1FJ81XAISort-Seq / Flow-Seq MPRA

NEB5α Pfrm library sort-seq, 100 µM formaldehyde

Sort-Seq Approach to Engineering a Formaldehyde-Inducible Promoter for Dynamically Regulated Escherichia coli Growth on Methanol

A randomly mutagenized 200-bp Escherichia coli Pfrm promoter library driving GFP on a plasmid was FACS-sorted in NEB5α after 100 µM formaldehyde induction. Public amplicon sequencing from expression bins 2–9 was aggregated into per-position/base enrichment versus the unsorted library.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

+100 µM formaldehyde

Plasmid Pfrm-GFP-Plac-FrmR reporter with an error-prone-PCR library across the native 200-bp Pfrm promoter; eight FACS expression bins were sequenced with Illumina MiSeq 2×201-nt amplicon reads. The public deposition contains sorted amplicon reads rather than barcode-linked RNA/DNA matrices, so the processed table reports aggregate nucleotide enrichment at each promoter position.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (23 of 23)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 23 definitions
condition
Resolved biological condition for this table
sort_bin
Deposited FACS expression-bin label; this experiment uses bins 2 through 9
sort_bin_rank
Numeric ordering of the FACS bin
promoter_position
1-based position within the 200-bp native Pfrm reference
variant_id
Position/base substitution key in the form Pfrm_posNNN_reference>observed
reference_base
Native base in the E. coli Pfrm reference
observed_base
Base observed at this position in retained sorted amplicon reads
is_reference_base
1 when observed_base equals reference_base, otherwise 0
forward_base_count
Deduplicated retained forward-orientation sequence count for this base and position
reverse_base_count
Deduplicated retained reverse-orientation sequence count after reverse-complement normalization
base_count
Sum of forward and reverse deduplicated counts
valid_forward_reads
All deduplicated retained forward sequences covering this promoter position
valid_reverse_reads
All deduplicated retained reverse sequences covering this promoter position
valid_reads
Total deduplicated retained sorted sequences covering this promoter position
unsorted_forward_base_count
Deduplicated forward count for this base and position in the unsorted-library baseline
unsorted_reverse_base_count
Deduplicated reverse count for this base and position in the unsorted-library baseline
unsorted_base_count
Sum of unsorted forward and reverse counts
unsorted_valid_forward_reads
All deduplicated retained unsorted forward sequences covering this position
unsorted_valid_reverse_reads
All deduplicated retained unsorted reverse sequences covering this position
unsorted_valid_reads
Total deduplicated retained unsorted sequences covering this position
frequency_within_bin
Raw frequency of observed_base among deduplicated retained sequences in this sorted bin
unsorted_frequency
Raw frequency of observed_base among deduplicated retained sequences in the unsorted library
log2_enrichment_vs_unsorted
Log2 of smoothed sorted-bin frequency divided by smoothed unsorted frequency, using a 0.5-count pseudocount

Quality control

Applied fixed-length read QC by retaining exactly 201-nt reads with the expected forward or reverse primer anchor, only A/C/G/T promoter sequence, and Hamming distance ≤17 from the native Pfrm span, following the paper's approximate distance filter. Exact duplicate promoter fragments were then removed within each sorted population and the unsorted baseline, as in the paper's redundant-sequence filter. Forward and reverse reads were orientation-normalized and aggregated; rows were retained only when both the sorted-bin and unsorted baseline had at least 50 covering sequences. Position 1 was omitted because it has no retained public-read coverage.

Curation notes

The public SRA files are aggregate sort-seq amplicon reads, not individual barcode-linked constructs or RNA/DNA activity ratios. Exact duplicate trimmed promoter fragments were removed within each population before counting. Forward reads cover native Pfrm positions 2–157 and reverse reads cover positions 43–200 after primer anchoring and orientation normalization; position 1 is therefore absent. The deposited NEB5α bin labels are retained as 2–9. The paper's sequence-analysis figures describe the same eight-bin experiment with a different low/high bin numbering convention. The common unsorted-library baseline is included in every row for direct comparison.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.