RMH
80.1K views | +37 today
 
Scooped by mhryu@live.com
onto RMH
Today, 1:12 AM
Scoop.it!

Property guidance for protein sequence generative models with ProteinGuide | Nbt

Property guidance for protein sequence generative models with ProteinGuide | Nbt | RMH | Scoop.it

No principled framework exists for conditioning sequence generative models for protein engineering on auxiliary information, such as experimental data, without additional training of a generative model. Here we present ProteinGuide, a method for such ‘on-the-fly’ conditioning. ProteinGuide is amenable to a broad class of protein generative models including masked language models such as ESM3, any-order autoregressive models such as ProteinMPNN and diffusion and flow-matching models on discrete state-spaces such as MultiFlow. ProteinGuide stems from a unifying statistical framework for these model classes. As proof of principle, pretrained generative models are used to design proteins with user-specified properties, such as higher stability or activity. Proteins are additionally designed to optimize for two desired properties that are in tension with each other. Lastly, we apply ProteinGuide jointly with wet-lab data generation to increase the editing activity of an adenine base editor in vivo, resulting in a base editor with higher editing efficiency than was previously achieved using seven rounds of directed evolution. On-the-fly conditioning of pretrained protein generative models guides protein generation toward specific properties.

mhryu@live.com's insight:

3st, workflow 1. Start with pretrained generative model (already trained). 2. Get/generate experimental property data for some sequences (e.g., synthesize and assay ~1,000–2,000 variants). 3. Train a lightweight predictive model on that data 4. At generation time, run the pretrained model's normal sampling procedure, but at each step, reweight the next-token probabilities using the predictive model's gradient/likelihood signal (via the Bayes rule) — this is the "guidance" step. Two flavors are offered: TAG (fast, approximate) or DEG (exact, slower). 5. Sample new sequences — these come out already biased toward the property one want. 6. (Optional, iterative) Assay the new sequences, get more data, retrain the predictor, guide again — repeating this loop (as they did for the base editor: round 1 → assay → train predictor → round 2 guided design).

library (any origin) + function data → train predictor → use predictor to guide ESM3's generation process → get new candidate sequences → synthesize/test those new candidates in the wet lab.

No comment yet.
RMH
Your new post is loading...
Scooped by mhryu@live.com
Today, 1:56 AM
Scoop.it!

Engineering viral protease-operated nanobodies for programmable and orthogonal control of protein function | Ncm

Engineering viral protease-operated nanobodies for programmable and orthogonal control of protein function | Ncm | RMH | Scoop.it

Precise control of protein function in living cells is essential for engineering programmable biological systems and therapeutic applications. However, most current strategies act indirectly by altering protein stability, localization, or proximity rather than directly modulating binding activity. Here we present VIPbodies, nanobodies engineered with self-cleaving viral proteases that convert protease inhibition into drug-dependent antigen recognition. Using anti-mCherry nanobodies as prototypes, we create variants responsive to orthogonal viral protease–inhibitor pairs and extend the design to diverse nanobody scaffolds. Each VIPbody functions independently within the same cell, enabling multiplexed regulation of distinct targets. When incorporated into transcriptional circuits, VIPbodies mediate drug-tunable gene expression and execute all six canonical Boolean logic operations, providing a compact framework for programmable cellular computation. Beyond gene regulation, VIPbody circuits mediate bidirectional control of pyroptosis and selectively activate apoptotic or pyroptotic programs via caspase coupling, thereby enabling chemogenetic control of protein function and cell fate. Precise control of protein function in living cells is essential for engineering programmable biological systems. Here the authors present VIPbodies, nanobodies engineered with self-cleaving viral proteases that convert protease inhibition into drug-dependent antigen recognition.

mhryu@live.com's insight:

gene exp control, mode of regulation, 1str,  Nanobody with protease inserted in a flexible loop → protease active by default → cuts nanobody in two → misfolded, can't bind target; add matching antiviral drug → protease blocked → nanobody stays intact/folded → binds target normally.

No comment yet.
Scooped by mhryu@live.com
Today, 1:38 AM
Scoop.it!

Quantifying protein unfolding kinetics with a high-throughput microfluidic platform | csys

Quantifying protein unfolding kinetics with a high-throughput microfluidic platform | csys | RMH | Scoop.it
Even after folding, proteins sample unfolded intermediates at risk of irreversible alteration (e.g., via proteolysis, aggregation, or posttranslational modification). Thus, kinetic stability impacts protein lifetime and abundance. However, we have very few measurements of unfolding rates, largely due to technical challenges. To address this, we developed SPARKfold (simultaneous proteolysis assay revealing kinetics of folding), a microfluidic platform to express, purify, and measure unfolding rate constants at high throughput via native proteolysis. We applied SPARKfold to determine unfolding rate constants for 1,104 protein samples comprising 31 dihydrofolate reductase orthologs with up to 78 chamber replicates each, providing statistical power to resolve subtle effects. SPARKfold rate constants for 5 constructs agreed with traditional measurements across a 150-fold range and provided information about the folding transition state via φ analysis. In future work, SPARKfold can reveal mutations that drive misfolding and aggregation and enable the rational design of kinetically hyperstable variants for industrial use. 
mhryu@live.com's insight:

fordyce, hts, N-terminal SNAP tag – protein – C-terminal EGFP tag, joined by flexible linkers. SNAP tag allows covalent surface attachment. GFP fluorescence is imaged to measure how much full-length protein was expressed/immobilized per chamber. Protease challenge → Thermolysin (a protease that only cuts exposed, unfolded regions — not the folded native state) is flowed into all chambers at once. Time-course imaging → GFP fluorescence is imaged repeatedly over time (0 to ~10,000 s) → as protein unfolds transiently, thermolysin cleaves it, releasing/destroying the GFP signal → fluorescence decays faster for proteins that unfold more readily.

No comment yet.
Scooped by mhryu@live.com
Today, 1:30 AM
Scoop.it!

Artificial intelligence catalyzes antimicrobial peptide design | ssb

Artificial intelligence catalyzes antimicrobial peptide design | ssb | RMH | Scoop.it
With broad-spectrum, low resistance, and multifunctional properties, antimicrobial peptides (AMP) are promising therapeutic agents against drug-resistant pathogens, yet their discovery and optimization still remain challenging due to the complexity of sequence-function associations. Artificial intelligence (AI), through the construction of comprehensive data-driven models that assisted with miscellaneous learning strategies, enables de novo peptide design by learning latent representations inherent in peptide sequences as well as their biological properties to ensure physically plausible and biologically relevant predictions. Consequently, this paradigm enhances the likelihood of designing peptide candidates with significantly improved therapeutic potential, reducing resource-intensive trial-and-error processes and revealing the transformative impact of computational innovation in advancing next-generation therapeutics. Here, we provide a snapshot of this field and survey two modes of AI-driven technologies for AMP design, one concentrated on identifying whether current data possess antimicrobial activity (identification-oriented) and the other on generating AMP candidates with potential therapeutic properties (generation-oriented). We also highlight the challenges and limitations that still hinder AMP development even accelerated by AI, as well as the foreseeable prospects, from finer-grained explorations to model-driven data enrichment and model enhancement.
No comment yet.
Scooped by mhryu@live.com
Today, 1:23 AM
Scoop.it!

Phage hijacks host phosphorothioate DNA modification machinery to circumvent bacterial Ssp defences | Nmb

Phage hijacks host phosphorothioate DNA modification machinery to circumvent bacterial Ssp defences | Nmb | RMH | Scoop.it

The bacterial Ssp defence system discriminates self from non-self by introducing sequence-specific phosphorothioate (PT) modifications in host DNA (via SspABCD) and cleaving unmodified foreign DNA (via SspFGH or SspE). Here we report PptA, a phage-encoded [4Fe–4S] cluster-containing protein, which hijacks cognate host cysteine desulfurase IscS homologues to assemble a streamlined PT modification machinery. Integrated biochemical and structural data delineate a model for intermolecular sulfur transfer within the IscS–PptA complex. Upon infection, robust expression of PptA, not merely its presence, drives sufficient PT incorporation into the phage genome, enabling molecular mimicry of host PT patterns. By masquerading as ‘self’, the modified phage DNA evades recognition and cleavage by SspFGH/SspE. Notably, PptA can reprogramme the Ssp-sensitive λ phage into an immune-evasive variant. These results reveal a co-evolutionary strategy used by phages to overcome PT-based bacterial immunity and provide a foundation for engineering therapeutic phages that bypass this widespread defence system. The phage-encoded protein PptA enables its DNA to masquerade as the host and evade cleavage, providing a potential foundation for engineering therapeutic phages that could bypass this widespread defence system.

No comment yet.
Scooped by mhryu@live.com
Today, 12:43 AM
Scoop.it!

Quorum sensing-driven metabolic altruism of nitrite-oxidizing bacteria fuels nitritation | Nwt

Quorum sensing-driven metabolic altruism of nitrite-oxidizing bacteria fuels nitritation | Nwt | RMH | Scoop.it

Nitritation, the conversion of ammonia to nitrite without further oxidation, offers an energy-efficient route for nitrogen removal, but its application is limited by the difficulty of selectively suppressing nitrite-oxidizing bacteria (NOB). The underlying biological mechanisms that enable such suppression remain poorly understood. Here we show that quorum sensing (QS), a cell–cell communication system, enables nitritation by regulating NOB behaviour. Using multi-omics and single-cell Raman spectroscopy, we demonstrate that QS signalling induces the overexpression of nirB in the dominant NOB genus Nitrospira, triggering an altruistic nitrite reduction causing self-inactivation. In contrast, ammonia-oxidizing bacteria refrain from this altruistic metabolism, gaining a decisive competitive advantage and directing nitrification flux towards nitritation. QS manipulation confirms that active QS is required to maintain nitritation, and single-cell analysis reveals that QS drives a stress-tolerant Nitrospira cell into a susceptible state, markedly reducing survival. These findings uncover an unknown social behaviour in the nitrifier community and offer new insights for nitritation stabilization. Nitritation is limited by the difficulty of suppressing nitrite-oxidizing bacteria, the mechanisms of which are clear. This study shows that quorum sensing drives altruistic self-inactivation in Nitrospira, favouring ammonia oxidizers and stabilizing nitritation.

mhryu@live.com's insight:

wastewater, nitrogen removal

No comment yet.
Scooped by mhryu@live.com
July 29, 11:49 PM
Scoop.it!

Evolution of a core ribosomal innovation in octopus | curB

Evolution of a core ribosomal innovation in octopus | curB | RMH | Scoop.it
Much of biology focuses on how genetic changes mediate new functions, but less attention is given to adaptations within the ancient molecular machines that execute the central dogma. Octopuses exhibit complex nervous systems and sophisticated behaviors that rival vertebrates but via an entirely divergent evolutionary history. Here, we serendipitously discovered that octopus ribosomes contain a structural break in the core ribosomal RNA that is unique among all animals. This break site enhances translation fidelity to reduce miscoding and subsequent protein aggregation, even when engineered into evolutionarily distant bacterial ribosomes. Furthermore, high-fidelity translation by octopus ribosomes supports proteomic stability during extensive RNA editing observed in cephalopods, suggesting synergy between distinct non-canonical modes of gene regulation. This adaptation emerged in recently derived octopuses with expanded nervous systems, thereby revealing a mechanism that could broadly support the evolution of novel organismal traits.
mhryu@live.com's insight:

2st, increase ribosome fidelity, 28S rRNA (large ribosomal subunit) is cleaved at one internal site into two separate pieces that stay non-covalently associated and still fold into a working ribosome. The break sits in helix 88 (H88), part of the tRNA-exit (E) site 

recreated it in E. coli inserted rRNA processing-stem sequences into the E. coli 23S rRNA gene at the matching position. in vitro, using the iSAT cell-free system: ribosomes were transcribed and assembled from scratch in a test tube.

fidelity measured: A luciferase reporter with a catalytic-site mutation (K529) that only produces active enzyme if the ribosome makes a translation error at that codon. Comparing luminescence from mutant vs. WT reporter mRNA gives an error-rate readout.

No comment yet.
Scooped by mhryu@live.com
July 29, 10:58 PM
Scoop.it!

gaftools: a toolkit for analyzing and manipulating pangenome alignments | bft

gaftools: a toolkit for analyzing and manipulating pangenome alignments | bft | RMH | Scoop.it

Linear reference genomes are ubiquitously used in genomics research, despite known biases associated with their use. In recent years, there has been a shift towards graph-based reference genomes to address some of these biases, which has required development of new algorithms and file formats. This has created a necessity for new tools capable of utilizing these formats and performing operations similar to those carried out by traditional methods. In this paper we present “gaftools”, a multi-purpose tool that introduces several utilities for processing graph alignments in GAF format. Gaftools enables users to index and sort alignments, with graph ordering serving as a necessary step for the sorting process. Additionally, it allows users to view subsets of alignments and perform realignment using the wavefront alignment algorithm, among other features. Many of these functionalities are inspired by SAMtools, which provides similar operations for linear genomes, while gaftools adapts and extends them for pangenomes.

No comment yet.
Scooped by mhryu@live.com
July 29, 3:03 PM
Scoop.it!

Design Principles of Plant Thermotolerance: Integrating Photosynthesis, Reproduction, and Field Resilience in a Warming Climate | pce

Design Principles of Plant Thermotolerance: Integrating Photosynthesis, Reproduction, and Field Resilience in a Warming Climate | pce | RMH | Scoop.it

Global food security is increasingly threatened by climate change, as rising temperatures compromise the yields of major staple crops, including wheat, rice, and maize. Enhancing plant thermotolerance has therefore become a critical priority for sustaining agricultural productivity. However, plant heat-stress responses have often been described as fragmented and pathway-specific, limiting their translation into effective crop improvement strategies. Here, we synthesize current knowledge of plant responses to heat stress and reframe them as an integrated set of design principles centered on preserving photosynthetic carbon gain under high temperature. By doing so, we provide a collection of insights that may help guide future research efforts and the development of strategies for improving plant thermotolerance. We organize major defense strategies into six functional domains: (i) membrane systems and structural integrity, (ii) photosynthetic regulation, (iii) protective metabolites and hormonal signaling, (iv) reactive oxygen species (ROS) scavenging, (v) protein homeostasis, and (vi) transcriptional and post-transcriptional regulation. Rather than treating these responses as independent pathways, we emphasize their temporal hierarchy, energetic costs, and functional interconnections, highlighting their shared objective—maintaining CO2 assimilation, energy balance, and biomass accumulation as thermal damage accelerates. We further discuss how insights from mutagenesis, transgenic approaches, and targeted genetic modification can be translated into crop improvement, clarifying opportunities and trade-offs that emerge when thermotolerance is engineered at distinct physiological nodes. Together, this design-centered framework provides a unifying conceptual and practical roadmap for developing high-yielding, heat-tolerant cultivars, offering actionable guidance for sustaining crop productivity in a warming world.

No comment yet.
Scooped by mhryu@live.com
July 29, 2:59 PM
Scoop.it!

Coevolution-driven reconstruction of multi-taxa siderophore interaction networks reveals topological diversity of microbial exploitation | brveco

Coevolution-driven reconstruction of multi-taxa siderophore interaction networks reveals topological diversity of microbial exploitation | brveco | RMH | Scoop.it

Microbial communities are shaped by secreted metabolites that mediate ecological interactions, yet predicting these interactions from genomic sequences remains difficult, because the specific recognition between co-functional metabolites (CFMs), such as siderophores, and their receptor proteins (Rec) cannot be inferred from gene annotation alone. This difficulty arises from three factors: the prevalence of Rec-mediated exploitation, the lack of high-accuracy functional annotations, and the absence of genomic co-localization between functionally paired CFM-Rec in Gram-positive bacteria. Here we present the Coevolution-based Interaction Model (CIM), an automated framework that maps specific CFM-Rec pairings directly from uncurated genomic datasets. Using a dynamic joint optimization strategy that accounts for exploitation asymmetry and avoids combinatorial explosion, CIM identifies functional pairings solely through evolutionary covariation. We validated this approach by reconstructing macroscale iron scavenging networks across nine bacterial taxa. Experiments confirmed that CIM can bridge genomic distances exceeding 3 Mb in the Gram-positive genus Rhodococcus to identify unlinked cognate receptors, and can accurately predict cross-utilization by exploiter strains despite substantial receptor sequence heterogeneity in Burkholderiaceae and Rhizobiaceae. Finally, topological analysis of the reconstructed networks shows that siderophore exploitation acts as a universal topological glue, fusing fragmented microbial populations into highly connected communities, and that the exploitability of siderophore production reverses depending on network modularity. CIM thus offers a scalable, sequence-to-ecology approach for predicting interactions mediated by secondary metabolites in microbial communities.

mhryu@live.com's insight:

1str

No comment yet.
Scooped by mhryu@live.com
July 29, 2:52 PM
Scoop.it!

Mol-CADiff: text-conditional molecule generation via causality-aware autoregressive diffusion | Ncm

Mol-CADiff: text-conditional molecule generation via causality-aware autoregressive diffusion | Ncm | RMH | Scoop.it

The design of molecules with desired properties is a key challenge in drug discovery and materials science. Traditional methods rely on trial-and-error, while recent deep-learning approaches accelerate molecular generation. However, existing models struggle with generating molecules based on specific textual descriptions. We introduce Mol-CADiff, a diffusion-based framework that uses causal attention mechanisms for text-conditional molecular generation. Our approach explicitly models the causal relationship between textual prompts and molecular structures, overcoming limitations in existing methods. We enhance dependency modeling both within and across modalities, enabling precise control over the generation process. While primarily designed for text-guided tasks, this architecture inherently supports unconditional generation, providing the added capability to autonomously sample the broader chemical space without explicit constraints. Here we show that Mol-CADiff outperforms alternative methods in generating diverse, chemically valid molecules, with better alignment to specified properties, enabling more intuitive language-driven molecular design. By bridging these modalities, our framework provides a versatile method for drug discovery. Computational approaches to molecular design often explore only limited regions of the vast chemical space. This study presents a causality-aware diffusion model that generates valid and diverse molecules with or without text prompts, improving controllability in molecular design.

No comment yet.
Scooped by mhryu@live.com
July 29, 2:46 PM
Scoop.it!

Sequential reading of a stepwise-shortened peptide immobilized on nanopore | nat

Sequential reading of a stepwise-shortened peptide immobilized on nanopore | nat | RMH | Scoop.it

Accurate decoding of peptide sequences is crucial in proteomics. However, achieving this goal is a technical challenge, owing to the compositional and structural complexity of peptides. Studies inspired by the success of nanopore nucleic acid sequencing have shown that nanopore-based techniques can also be applied to peptide sequencing. The key is to generate narrowly distributed, consistent and sequence-dependent events during sequential nanopore readout. Here we introduce a nanopore-based strategy termed transient pore analyte looping (tPAL). We develop an engineered Mycobacterium smegmatis porin A (MspA) nanopore that is dual modified with a nickel-ion-bound nitrilotriacetic acid (NTA-Ni) adapter and the target peptide. This distinctive sensing configuration enables precise recognition of the N terminus of the immobilized peptide by multiple re-readings. With the aid of cholesterolized aminopeptidase, the immobilized peptide can be shortened sequentially in single-amino-acid increments, yielding sequence-dependent, stepwise and narrowly distributed signal alterations that provide clues to allow peptide sequence decoding. Our strategy achieves single-amino-acid resolution and effectively identifies single-amino-acid mutations, post-translational modifications and unnatural-amino-acid insertions—indicative of its versatility in nanopore proteomics and chiral peptide analyses. A nanopore-based ‘chop and measure’ method sequences peptides at single-amino-acid resolution by using enzymatic digestion to progressively shorten the N terminus one residue at a time, together with repetitive N-terminus re-reading.

mhryu@live.com's insight:

The target peptide (engineered with a C-terminal cysteine) is attached to the nanopore once, at the start. The nanopore itself carries two handles on different monomers: an NTA-Ni adapter and a maleimide-PEG4 linker. When the peptide is added to the chamber, its free N-terminus is transiently grabbed by NTA-Ni → this positions the C-terminal Cys next to the maleimide group → covalent conjugation locks the peptide permanently onto the pore as peptide–PEG4–MspA–NTA-Ni.

From there the sequencing runs as a repeating cycle rather than a fresh loading step each time: exposed N-terminus reversibly binds NTA-Ni (generating many repeated current reads of that residue, since binding/release keeps happening even after covalent tethering) → cholesterol-anchored aminopeptidase (sgAP-dA15-chol, enriched right at the membrane) clips off that exposed N-terminal residue the moment it's momentarily released from the adapter → the newly shortened peptide exposes a new N-terminus → that terminus is captured by NTA-Ni again, producing a new distinct signal stage (aa1 → aa2 → aa3 → ...) → cycle repeats, processively working from the N-terminus toward the fixed C-terminal anchor, until only the tethering cysteine is left.

No comment yet.
Scooped by mhryu@live.com
July 29, 1:23 AM
Scoop.it!

Understanding metabolic interactions between bacteria and fungi for cross-kingdom consortia applications | cin

Understanding metabolic interactions between bacteria and fungi for cross-kingdom consortia applications | cin | RMH | Scoop.it
Bacterial–fungal interactions represent fundamental ecological associations that shape microbial community structure across diverse environments. While traditionally framed through the lens of antagonism, bacteria and fungi engage in sophisticated metabolic dialogs extending far beyond simple warfare. Both primary and specialized metabolites function as context-dependent signals, nutrient resources, and modulators of cellular processes that fundamentally influence the physiology, development, and evolutionary trajectory of both fungi and bacteria. Primary metabolites mediate mutualistic relationships through cross-feeding and syntrophy, while specialized metabolites, including volatile organic compounds, lipopeptides, and phenazines, modulate fungal physiology at sub-inhibitory concentrations, reprogramming metabolic networks and triggering adaptive responses without causing cell death. At the molecular level, bacterial metabolites regulate fungal gene expression through transcriptional reprogramming, with consequences that extend to long-term evolutionary adaptation. The complexity of Bacillus–Trichoderma interactions exemplifies how these principles translate into ecological outcomes, exhibiting context-dependent transitions from competition to synergism that enhance biocontrol efficacy, plant growth promotion, and organic matter turnover. Recognizing the multifunctional nature of microbial metabolites beyond direct toxicity opens new avenues for rationally engineering cross-kingdom consortia with targeted agricultural and biotechnological applications.
No comment yet.
Scooped by mhryu@live.com
July 29, 1:19 AM
Scoop.it!

EU relaxes rules for gene-edited crops | Nbt

EU relaxes rules for gene-edited crops | Nbt | RMH | Scoop.it

The new laws break away from 20-year-old restrictive GMO directives to give an official nod to plants made with new genomic techniques.

mhryu@live.com's insight:

The EU legislation creates two categories of NGT plants, defined by the genetic alterations they undergo. The NGT-1 designation bypasses GMO rules; it covers gene-edited plants that could theoretically arise naturally or through conventional breeding techniques. This includes plants generated by targeted mutagenesis, which include a maximum of 20 nucleotide insertions or substitutions, as well an unlimited number of deletions. For each protein-coding sequence, just three genetic modifications are permitted, but non-coding regions, including introns and regulatory elements, are not subject to that limit. The NGT-1 designation also applies to plants obtained by cisgenesis, which involves DNA insertions or substitutions involving genetic material from the plant’s natural breeding pool, as well as inversions or translocations of endogenous sequences. None of these constructs are permitted to cause interruptions of endogenous genes or the creation of new gene sequences encoding chimeric proteins. A total of 20 genetic modifications across these two subcategories is permitted for each genome copy within a particular plant species.

No comment yet.
Scooped by mhryu@live.com
Today, 1:44 AM
Scoop.it!

PlasChain: an algorithm for improving long plasmid reconstruction from metagenome assemblies | brvbi

Plasmids play a critical role in horizontal gene transfer and the spread of antibiotic resistance. However, recovering complete plasmid sequences from metagenomic samples remains highly challenging due to extensive repeat content, structural heterogeneity, and large variation in plasmid size. Existing methods typically identify plasmids from metagenome assemblies by exploiting coverage differences or detecting minimum-weight cycles in assembly graphs. While effective for dominant plasmids, these approaches often fail to recover low-abundance and long plasmids.  Here we present PlasChain, a novel algorithm designed to improve plasmid assembly and identification from complex metagenomic data. Building upon the cycle-peeling strategy of SCAPP, PlasChain incorporates contig path information and a cycle-merging procedure to prevent long plasmids from being fragmented into multiple shorter cycles. In addition, PlasChain jointly leverages paired-end read alignments, sequence composition patterns, and coverage variation to filter out assembly artifacts and reduce false positives. We evaluated PlasChain against state-of-the-art plasmid assemblers, including SCAPP and metaplasmidSPAdes, using a diverse set of simulated and real metagenomic datasets. Across nearly all benchmarks, PlasChain demonstrates superior performance in recovering long plasmids while maintaining competitive accuracy in assembling short plasmids. Furthermore, analysis of real metagenomic samples shows that PlasChain is capable of assembling previously uncharacterized plasmids, including putative megaplasmids that are typically underrepresented in current plasmid databases. PlasChain is a novel graph-based plasmid assembler that improves the recovery of long plasmids from short-read metagenomic data. Evaluation on diverse simulated and real metagenomic datasets demonstrates that PlasChain consistently outperforms existing plasmid assemblers, especially for long plasmid reconstruction. These results highlight the potential of PlasChain to facilitate comprehensive characterization of plasmid diversity, antimicrobial resistance, and horizontal gene transfer in complex microbial communities. 

No comment yet.
Scooped by mhryu@live.com
Today, 1:33 AM
Scoop.it!

PlantMDCS: A code-free, modular toolkit for rapid deployment of plant multi-omics databases | pcm

PlantMDCS: A code-free, modular toolkit for rapid deployment of plant multi-omics databases | pcm | RMH | Scoop.it
The rapid accumulation of plant multi-omics datasets has increased the need for systems that support local data management, analysis, and reuse. However, conventional web-based databases construction often requires programming expertise and long-term maintenance, limiting its accessibility for small research groups and project-specific datasets. Here, we developed the Plant Multi-omics Database Construction System (PlantMDCS), a locally deployable, user-friendly graphical platform for constructing and managing plant multi-omics databases. PlantMDCS adopts a decoupled front-end and back-end architecture. The back end serves as the core engine for data management and computation, and is responsible for the storage, integration, and hierarchical association of multi-omics data. Once initialized, the front end supports the complete research workflow, including data import, querying, integrative analysis and visualization. All operations can be performed without programming, while local resource usage is mainly dominated by the disk storage required for user-provided datasets rather than sustained computational overhead. Benchmarking across representative plant datasets, including alfalfa, maize, rice, poplar, and wheat, showed that PlantMDCS can efficiently construct local database modules from appropriately prepared datasets. Concurrent-access testing defined the stable operating range and scalability limits under local deployment. A publicly accessible PlantMDCS-MOD database further demonstrates deployment of PlantMDCS-generated multi-species and multi-omics database entries. An end-to-end alfalfa transcriptome–metabolome case study further demonstrated the biological applicability of PlantMDCS. Together, PlantMDCS provides a practical framework for transforming fragmented file-based plant multi-omics workflows into reusable, locally controlled, and database-supported analytical systems.
No comment yet.
Scooped by mhryu@live.com
Today, 1:29 AM
Scoop.it!

Design principles for molecular Raman sensors: A sea of opportunities | cin

Design principles for molecular Raman sensors: A sea of opportunities | cin | RMH | Scoop.it
Molecular Raman sensors come with an unprecedented opportunity and promise of allowing simultaneous imaging and tracking of multiple bioanalytes in living cells. How can we achieve this dream goal for molecular imaging? Careful and strategic design of Raman sensors will be key. While we can utilize insights from foundational principles underlying molecular fluorescence sensor design to acheive selective detection, handling the inherent low sensitivity of Raman scattering is vital. Along with instrumentation, basic organic chemistry and molecular spectroscopy concepts can be leveraged to improve the sensitivity of Raman sensors to achieve live-cell imaging of endogenous levels of bioanalytes. In this Current Opinion paper, we attempt a ‘Molecular Raman sensor design 101’ for the chemical biologist, highlighting lessons from recent exciting examples of responsive Raman sensors.
No comment yet.
Scooped by mhryu@live.com
Today, 1:12 AM
Scoop.it!

Property guidance for protein sequence generative models with ProteinGuide | Nbt

Property guidance for protein sequence generative models with ProteinGuide | Nbt | RMH | Scoop.it

No principled framework exists for conditioning sequence generative models for protein engineering on auxiliary information, such as experimental data, without additional training of a generative model. Here we present ProteinGuide, a method for such ‘on-the-fly’ conditioning. ProteinGuide is amenable to a broad class of protein generative models including masked language models such as ESM3, any-order autoregressive models such as ProteinMPNN and diffusion and flow-matching models on discrete state-spaces such as MultiFlow. ProteinGuide stems from a unifying statistical framework for these model classes. As proof of principle, pretrained generative models are used to design proteins with user-specified properties, such as higher stability or activity. Proteins are additionally designed to optimize for two desired properties that are in tension with each other. Lastly, we apply ProteinGuide jointly with wet-lab data generation to increase the editing activity of an adenine base editor in vivo, resulting in a base editor with higher editing efficiency than was previously achieved using seven rounds of directed evolution. On-the-fly conditioning of pretrained protein generative models guides protein generation toward specific properties.

mhryu@live.com's insight:

3st, workflow 1. Start with pretrained generative model (already trained). 2. Get/generate experimental property data for some sequences (e.g., synthesize and assay ~1,000–2,000 variants). 3. Train a lightweight predictive model on that data 4. At generation time, run the pretrained model's normal sampling procedure, but at each step, reweight the next-token probabilities using the predictive model's gradient/likelihood signal (via the Bayes rule) — this is the "guidance" step. Two flavors are offered: TAG (fast, approximate) or DEG (exact, slower). 5. Sample new sequences — these come out already biased toward the property one want. 6. (Optional, iterative) Assay the new sequences, get more data, retrain the predictor, guide again — repeating this loop (as they did for the base editor: round 1 → assay → train predictor → round 2 guided design).

library (any origin) + function data → train predictor → use predictor to guide ESM3's generation process → get new candidate sequences → synthesize/test those new candidates in the wet lab.

No comment yet.
Scooped by mhryu@live.com
Today, 12:38 AM
Scoop.it!

Rational design of disordered proteins for sequence–function investigation | nat 

Rational design of disordered proteins for sequence–function investigation | nat  | RMH | Scoop.it

Despite lacking a stable three-dimensional structure, intrinsically disordered protein regions (IDRs) are ubiquitous across all kingdoms of life and have essential cellular roles. While rational design of folded proteins has seen substantial recent progress, our ability to design IDRs remains more limited. Here we present GOOSE (Generate disOrdered prOteins Specifying propErties), a comprehensive computational framework for the rational design of IDRs. GOOSE’s versatility and throughput enable us to design and test thousands of IDR sequences to reveal distinct sequence-to-function relationships. Using GOOSE to explore these relationships, we examine how sequence properties influence IDR structural ensembles in cells, design IDRs that respond to structural changes associated with cell volume decrease, create scaffold IDRs that self-assemble and recruit specific clients, and design novel IDRs that protect cells from desiccation. Our work uses rational sequence design as a powerful method for exploring function in IDRs and provides a versatile tool for designing functional disordered proteins. GOOSE enables the design and testing of thousands of disordered protein region sequences to reveal distinct sequence-to-function relationships.

mhryu@live.com's insight:

2st, llps

No comment yet.
Scooped by mhryu@live.com
July 29, 11:28 PM
Scoop.it!

Carbon-to-ATP ratios across the kingdoms of life | iSci

Carbon-to-ATP ratios across the kingdoms of life | iSci | RMH | Scoop.it
Estimating biomass is essential for quantifying energy flow, nutrient cycling, and ecosystem functioning across the biosphere. For nearly 60 years, the cellular carbon-to-ATP ratio (C:ATP, g/g) has been used to estimate marine microbial biomass, often assuming a value near 250. Here we compile ∼400 measurements from more than 80 studies spanning bacteria, unicellular eukaryotes, animals, plants, and tissues, and show that C:ATP varies by five orders of magnitude across species, physiological states, and environmental conditions. We then develop a mechanistic model demonstrating this variation arises from differences in ATP production rate, ATP turnover time, and the fraction of metabolically active carbon. The model predicts that median C:ATP is within a factor of two of 250 for bacteria and unicellular marine eukaryotes, while systematically differing in other groups. These results indicate that C:ATP should be treated as a context-dependent physiological quantity in biomass estimation and metabolic modeling across organisms and environments.
No comment yet.
Scooped by mhryu@live.com
July 29, 10:45 PM
Scoop.it!

SynFlow: an interactive online genome structural variant viewer | bft

SynFlow: an interactive online genome structural variant viewer | bft | RMH | Scoop.it

Structural variations (SVs), including inversions, translocations, duplications, and large insertions or deletions, are key drivers of genome evolution and phenotypic diversity. With the increasing number of high-quality, chromosome-scale genome assemblies, the ability to detect and interpret SVs has become a crucial aspect of modern genomics. While SV detection has advanced, most visualization methods produce static plots that fall short when researchers, particularly in comparative genomics, need to interactively explore large datasets, zoom into specific genomic regions, or dynamically filter structural events in real time. To address this gap, we introduce SynFlow, a lightweight, web-based interactive application specifically designed for exploring and visualizing structural variations identified by SyRI. We demonstrate that SynFlow can reproduce complex static synteny plots published in literature, but transforms them into dynamic, shareable visualizations that support real-time filtering, reordering, and deep exploration of specific SVs, including translocations. SynFlow is available as a web server and offers multiple entry points: browsing precomputed datasets (e.g., banana and grapevine genomes), uploading user-provided SyRI outputs, or running an integrated workflow to produce and visualize SVs on the fly.

No comment yet.
Scooped by mhryu@live.com
July 29, 3:02 PM
Scoop.it!

Aquatic bacterial metabolic rates from RNA quantitative stable isotope probing (RNA-qSIP) depend on experimental design | brvt

Aquatic bacterial metabolic rates from RNA quantitative stable isotope probing (RNA-qSIP) depend on experimental design | brvt | RMH | Scoop.it

Quantitative stable isotope probing (qSIP) allows researchers to calculate taxon-specific carbon incorporation from sequencing of natural microbial communities, which can be used as a proxy for metabolic activity rates and subsequently as an input for biogeochemical modeling. While qSIP is widely utilized in soils to investigate the identity and metabolic activity of largely unculturable microbes, the application of qSIP in marine and aquatic ecosystems is more recent. Here, we investigated how bioreactor type (batch vs. chemostat) and carbon substrate complexity (single vs. multiple substrates) affect the incorporation of 13C-labeled glucose into rRNA after 24 hours using excess atomic fraction (EAF) as a proxy for metabolic activity rate. We found that the growth dynamics and community composition of the 13C-incorporating bacteria differed significantly for each treatment. EAF was positively correlated with both 16S gene copy number and a genomic index of copiotrophy in both batch treatments, but not in the chemostat, suggesting that chemostats dampen the competitive advantage of fast-growing copiotrophic taxa. Our results demonstrate that both substrate complexity and experimental regime influence qSIP-derived metabolic activity estimates and provide guidance for future applications of qSIP in aquatic environments.

No comment yet.
Scooped by mhryu@live.com
July 29, 2:54 PM
Scoop.it!

Structure-Guided Redesign of Terminal Deoxynucleotidyl Transferase Enables Scalable Enzymatic DNA Synthesis for Data Storage | angeC

Structure-Guided Redesign of Terminal Deoxynucleotidyl Transferase Enables Scalable Enzymatic DNA Synthesis for Data Storage | angeC | RMH | Scoop.it

DNA, with its exceptional information capacity and chemical stability, represents a promising material for next-generation data storage to meet the exponential growth of global digital information demands. Enzymatic DNA synthesis provides a sustainable route to DNA production. However, the practical scalability of the underlying polymerization chemistry has been fundamentally constrained by the low catalytic efficiency and aggregation-induced inactivation of terminal deoxynucleotidyl transferase (TdT) that is responsible for nucleotide polymerization. Here, we report a structure-guided enzyme design framework that overcomes these intrinsic limitations by decoupling solubility and catalytic performance in a processive polymerase. Computational redesign of aggregation-prone regions markedly enhances soluble expression, while targeted active-site engineering improves catalytic efficiency toward 3′-ONH2-dNTPs used in enzymatic DNA synthesis. The resulting TdT variant HL2-LKI achieves 3.4 g L−1 soluble expression in a 5 L fermenter without fusion tags and exhibits high polymerization efficiency (99.9%) and DNA writing fidelity (98.9%). This redesign reduces enzyme production costs to approximately $0.7 g−1, nearly seven orders of magnitude lower than the catalog price of commercially available TdT. This work establishes a generalizable strategy for transforming aggregation-limited enzymatic polymerization reactions into scalable and low-cost molecular manufacturing processes, thereby advancing the practical implementation of DNA as an information material.

No comment yet.
Scooped by mhryu@live.com
July 29, 2:49 PM
Scoop.it!

Beyond enzyme engineering: ordered enzyme immobilization drives enhanced productivity in vitro | frn

Beyond enzyme engineering: ordered enzyme immobilization drives enhanced productivity in vitro | frn | RMH | Scoop.it

Increasing environmental and economic pressures associated with global fossil fuel demand necessitate a shift toward sustainable fuel production. Production of second-generation biofuels, such as isobutanol, presents a promising opportunity; however, product toxicity limits in vivo production, motivating the development of optimized in vitro systems. As prior efforts have focused on enzyme engineering to improve titer, these systems remain constrained by diffusion-limited mass transfer. Here, we introduce an engineered, ordered cellulosome-based immobilization system that facilitates enhanced enzymatic productivity. Using keto-acid decarboxylase, alcohol dehydrogenase, and formate dehydrogenase as a model for the final steps of the isobutanol pathway, the system achieved a preliminary isobutanol titer of 5.92 g/L, a 78.4% yield, and an enzymatic productivity of 0.34 mL-1 h-1, representing marked improvement over previous cell-free approaches. This proof-of-concept approach introduces the importance of targeted immobilization alongside enzyme optimization and demonstrates ordered scaffoldin-mediated immobilization as a versatile platform approach for future in vitro biofuel and biochemical production.

No comment yet.
Scooped by mhryu@live.com
July 29, 1:37 AM
Scoop.it!

The response regulator BqsR/CarR controls Fe2+ acquisition in Pseudomonas aeruginosa | Ncm

The response regulator BqsR/CarR controls Fe2+ acquisition in Pseudomonas aeruginosa | Ncm | RMH | Scoop.it

Pseudomonas aeruginosa is a ubiquitous, Gram-negative bacterium that forms biofilms and is responsible for antibiotic-resistant hospital-acquired infections in humans. The P. aeruginosa BqsRS two-component system regulates biofilm formation and dispersal by sensing extracytoplasmic Fe2+, but the mechanistic details of this process are poorly understood. In this work, we report the crystal and solution structures of the PaBqsR response regulator receiver domain, comprising a (βα)5 response regulator assembly, and the DNA-binding domain, comprising a helix-turn-helix motif. Consistent with its cognate stimulus being Fe2+, we show that BqsR binds directly to the promoter region of the feo operon that encodes the bacterial Fe2+ transport system FeoABC. Corroborating these in vitro results, transcriptional studies show that BqsR is a global regulator controlling many important genes in PAO1, including the feo operon. Intriguingly, promoter-based assays reveal that BqsR is a dynamic regulator that responds to bioavailable Fe2+, likely through the ability of BqsR to bind Fe2+ directly via a His-rich motif, independent of the BqsS membrane His kinase. To our knowledge, this mode of regulation has not been reported previously among OmpR-like response regulators but represents an important level of control over Fe2+ acquisition in P. aeruginosa that could be an attractive therapeutic target to treat hospital-acquired infections. The BqsRS two-component system regulates biofilms in the infectious pathogen Pseudomonas aeruginosa. Here, the authors show that the global response regulator BqsR controls Fe2+ acquisition in PAO1 through a previously unrecognized regulatory mechanism.

mhryu@live.com's insight:

iron sensor

No comment yet.
Scooped by mhryu@live.com
July 29, 1:21 AM
Scoop.it!

AI proteomics: from protein identification to virtual cells | Nmet

AI proteomics: from protein identification to virtual cells | Nmet | RMH | Scoop.it

Artificial intelligence (AI) is transforming scientific research, including proteomics. In this Perspective, we highlight key mass spectrometry (MS)-based proteomics areas where AI is driving innovation, ranging from protein identification to building AI virtual cells. These include improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and, ultimately, enabling AI virtual cells. Finally, we call for global collaboration among data producers, data consumers and other stakeholders to establish an AI-friendly ecosystem for MS-based proteomics, laying the foundation for transformative advancements in proteomics driven by AI. This Perspective highlights key research areas within mass spectrometry-based proteomics where AI is poised to drive significant advances.

No comment yet.