Each LLM call incurs latency, cost, and token overhead. More subtly, it compounds context:
every step includes not only the original query, but intermediate outputs and scratchpad logic from earlier prompts.
This creates a growing burden on both inference and model performance.
I was working with agents over a year ago before the common workflows had really been set in stone. At that time we were heavily doctoring the context to give a very streamlined representation of what had occurred during a given run to the LLM. Is this not standard practice?
From the article:
I was working with agents over a year ago before the common workflows had really been set in stone. At that time we were heavily doctoring the context to give a very streamlined representation of what had occurred during a given run to the LLM. Is this not standard practice?