Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This is super cool!

From the article:

  Each LLM call incurs latency, cost, and token overhead. More subtly, it compounds context:
  every step includes not only the original query, but intermediate outputs and scratchpad logic from earlier prompts. 
  This creates a growing burden on both inference and model performance.
I was working with agents over a year ago before the common workflows had really been set in stone. At that time we were heavily doctoring the context to give a very streamlined representation of what had occurred during a given run to the LLM. Is this not standard practice?


yes, AFAIK right now, there are no easy ways of "slimming" context because no one knows what it should be or how.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: