Papers
arxiv:2607.18144

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

Published on Jul 20
· Submitted by
Maksim Kuznetsov
on Jul 21
Authors:
,
,
,
,
,
,

Abstract

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

Community

Paper author Paper submitter

3D-Fit is a novel benchmark for evaluating the ability of large language models to generate 3D molecules under realistic spatial constraints for structure-based drug design. It tests whether general-purpose LLMs can produce ligands conditioned on a protein pocket while also satisfying additional requirements such as anchor fragments, pharmacophore points, and mandatory protein–ligand interactions. The benchmark uses token-efficient textual descriptions of 3D conditions and a structured output format for generated molecules.

The results show emerging 3D instruction-following: general-purpose LLMs can often satisfy explicit local constraints—especially coordinate-like anchors and pharmacophores—and can attempt multiple heterogeneous conditions in a single prompt. At the same time, satisfying those constraints is not enough for physically reliable binders. LLM poses frequently show steric clashes, weaker intramolecular validity, and worse docking scores than specialized diffusion models, even after local optimization.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.18144 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.18144 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.18144 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.