Home

FLAIR Lab · Mila & Université de Montréal

Machine learning for the language of life.

We develop methods across the language modeling pipeline for biological sequences such as proteins, genomes, and transcriptomes, with applications in drug discovery.

Recent news
Jul 30, 2026 New preprint: pLM representations unlock metagenomic space beyond homology. Pre-training 100 protein language models reveals how training data controls the trade-off between evolutionary calibration and structural modeling, enabling the retrieval and in silico validation of diverse enzyme candidates from billions of metagenomic sequences. With Lola Le Breton, David Heurtel-Depeiges, Douglas Millar, Lara Zetzsche, Robert Vernon, Christopher Langmead, and Sarath Chandar. In collaboration with Amgen.
Jul 23, 2026 New preprint: High-resolution dissection of concept acquisition in different families of protein language models. Layer-by-layer analysis of ESM2 and AMPLIFY maps a progression from physicochemical properties and motifs to secondary structure and domain-level semantics, revealing that data and compute shape concept emergence more than model size. With Shawn Whitfield, Tom Marty, Robert Vernon, Christopher Langmead, and Dhanya Sridhar. In collaboration with Amgen.
Jul 08, 2026 New preprint: A systematic analysis of machine learning pipelines for robust antimicrobial resistance prediction. Across nine clinically relevant species–antibiotic combinations, choices such as k-mer length can shift F1 scores by over 20 points, while tree-based models remain robust and interpretable. With Alex Aselstyne, Enamundram Naga Karthik, Meriem El Azami, Romain Pogorelcnik, and Sarath Chandar. In collaboration with bioMérieux.
Apr 23, 2026 Proud to see Lola Le Breton present NeoBERT: A Next Generation BERT at ICLR 2026 as part of the TMLR journal track. NeoBERT brings modern architecture, data, and pre-training to encoders, achieving state-of-the-art results on MTEB with just 250M parameters. With John X. Morris, Mariam El Mezouar, and Sarath Chandar.
Feb 27, 2026 New preprint: CoPeP: Benchmarking Continual Pretraining for Protein Language Models. Spanning a decade of UniProt updates and 31 protein tasks, CoPeP shows that continual learning can leverage temporal information to outperform naive pre-training at scale. With Darshan Patil, Pranshu Malviya, Mathieu Reymond, and Sarath Chandar. In collaboration with Genentech.

Open positions

We are recruiting one PhD student. Get in touch.