The effect of this is that it is encouraging longer agent threads. All of the previous models across major providers had a 10% cache read cost (vs normal cost) and not this is 5%
So longer threads get cheaper and one-shots stay the same price.
I've been spending all my time recently thinking about LLM costs. From that, I am currently of the mind that the most interesting question is if they can get it to work in the first place. After that, like with "performance" in the past, I believe there are ways to optimize things. We see the code harnesses doing this out of necessity, for example.
A blast from the past. When Taskrabbit was acquired by IKEA, I built several tools that went through the whole catalog via various crawling approaches. One tool was to estimate how long it would be to put each item together for an initial training set.
I am excited about the Rails defaults where background and cache and sockets are all database driven. For normal-sized projects that still need those things, it's a huge win in simplicity.
Ario | Onsite in Palo Alto, CA | Full-time | heyario.com
I recently started at Ario where we are building an app for parents that uses AI to make things easier. We are targeting saving them one hour every day.
Previously, I co-founded Taskrabbit. That was somewhat similar but now it's time to get LLMs in on the action!
It's in the app store, but still early in the game, so we are looking for engineers to build it out. We are using Python and React Native.
Unrelated (I'm based in Switzerland), but I wanted to say that from the perspective of a potential future user, your landing page is great at showcasing why parents could benefit from your app. Well done.
We have to work on the scaling. There is lots of web scraping and LLM calls that we need to make sure works under load. I'm sure there will be quality improvements as well.
Just came to say, I still think this is the best balance between the many factors of running a dev team. I keep trying to recreate it in every tool I use.
Airbyte acquired our Reverse ETL company, Grouparoo, 1.5 years ago. There is so much to solve making just the Extract and Load work well and so much value that comes from that, we have been busy there. I'm excited to circle back to publishing next year.
I like how the article notes that the stuff we were talking about with Reverse ETL (mostly activating your data in SaaS systems like Salesforce, Zendesk, etc) is one important part of Publishing. But we are also seeing traditional use cases like file uploads and new fancy stuff like vector databases.
I built a system for TaskRabbit that scraped all the IKEA products from a variety of sources and ran algorithms to determine their category and predict how long they would take to be assembled. Then there was a Mechanical Turk sort of system for human input. When combined with real-world feedback from the Taskers, it was pretty good.
For better or worse, I've personally been through the entire catalog multiple times.
Maybe as Turtle or JSON-LD, but you need a format that encapsulates triples and various data types. Otherwise you're throwing away most of the utility of a semantic web knowledge base.
So longer threads get cheaper and one-shots stay the same price.
reply