Constraint Drift in Large Language Models: Evaluating Reliability in Engineering Reasoning
Large language models (LLMs) are increasingly being used in engineering related contexts such as education, coding environments, and design assistance tools. While they are effective at generating explanations, solving structured problems, and supporting early stage ideation, their reliability in engineering reasoning tasks that involve multiple interacting constraints remains uncertain. Unlike general benchmark questions, engineering problems require constraints to remain consistent across multi-step reasoning processes rather than being applied only at the final step. Recent research in both general reasoning benchmarks and STEM-focused datasets suggests that performance decreases as task complexity and constraint interactions increase. This paper synthesizes findings from existing literature on both general LLM reasoning and STEM-specific evaluation benchmarks. Across these studies, a consistent pattern emerges: while models perform well on individual reasoning steps, their consistency declines when constraints must be maintained over longer reasoning chains. To describe this behavior, the concept of constraint drift is introduced, referring to the gradual loss of constraint consistency during multi-step reasoning. This framework is used to better understand limitations in current evaluation methods for engineering processes and applications.