Hacker Newsnew | past | comments | ask | show | jobs | submit | gvkhna's commentslogin

The harness is making a big difference, lackluster performance with pi but somehow very good performance with opencode. There’s some rl there for sure, for a smaller model it’s likely going to perform much better in a harness it understands the best.


I've been using a harness that I made myself, and making improvements to the harness has vastly improved its performance. It's actually been a useful model to identify flaws in the harness.

I think something also went wrong with the Ox provider last night (at least on OpenRouter), for a few hours it wouldn't accept tools. Zero change to the harness while I slept and it was back working again the next morning.


256gb ddr5, so this is going to cost north of 100k? Damn.


Seems like the memory is of a socketed type, so I'm sure they'd send you one without RAM if you want :D.


No way, it's only like 4x64 GB of ECC DDR5 5600 RDIMMs. Who's paying $100k for that? I'll gladly sell them this at a 10% discount.


If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?


Someone could've trained model on top of GLM. Same way Cognition trained their SWE model on top of Kimi and Cursor did same with their Composer model.


While possible the amount of variation in serving infrastructure is unlikely to land with actually giving the exact same errors zhipu does.

It feels like glm flash, and there was a report zhipu had secured a huge new cluster suggesting they have the capacity. My guess anyway.

https://www.tomshardware.com/tech-industry/artificial-intell...


The reasoning levels are the same as GLM 5.3. GLM 5.3 is still not open...

I believe it's GLM 5.3 Flash or Air.


Reasoning levels are often just injected system prompts so not a great way to fingerprint models.


But it's an error, not a response.


Ziphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5


There's three options here:

- The provider has a massive amount of (unused) hardware. Google or Cursor seem most likely

- The model is extremely efficient, beyond anything we've seen so far

- Whomever made the model has improved the cache efficiency in such a way that it's very cheap to serve. See e.g Deepseeks or Xiaomi caching (pre-price increase)


Option 4: the claimed capacity is not true. Real world usage hasn’t reached anywhere close to it.


    > 1 quadrillion tokens per day on Nous portal
If you are referring to this number (https://xcancel.com/NousResearch/status/2090899914700054780), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.


I was hitting 429 overloaded regularly with Ox on OpenRouter yesterday, but a lot of that turned out to be problems with my harness. I fixed some bugs, improved the back-off, and I haven't hit a 429 error since (touch wood).

OpenRouter says they're doing 6 Trillion tokens a day with Ox Alpha so far, and it has been their biggest launch of all time. OpenCode claimed they had capacity for 100T a day.

https://x.com/OpenRouter/status/2091912024922177562 https://x.com/opencode/status/2090544355824038300


That’s the whole point, just cost and compute limitations in your way (mostly).


Incidentally the fact that you could read and write from essentially thousands of harddrives at once for even a tiny file makes you wonder what novel use cases you could get out of that.


Nice @OP i put together something similar as well. Incidentally I found for motion design specifically llm is not able to infer specific animations as well as it just being described very plainly and accurately what is happening and the timing.

One thing which sort of worked decently was actually take the frames and put them into a grid and have the agent look at the image of all of the frames together. It did surprisingly well but missed a lot of subtle details that it couldn’t see.

Also tried various kinds of vision embeddings, heat map of motion etc, and blur etc to show motion. But none really worked as well so I ended up just describing it until it got it. Haven’t quite found the right solution yet.


Go is my go to for simple coding projects now, if i have an agent work grind on it. It’s lightweight fast and resource minimal, while being quick to code.

Go tooling is kind of surprisingly minimal and takes more taming then should be required. Especially monorepo workspace support seems a bit like an afterthought.

Hope to see continued improvements because in agentic era go is a serious go to.

Rust is top tier but for most things, it’s not really needed for an agent, to spend longer tokens fixing memory ceremony.

But rust tooling is best in class overall I feel and hope google continues improvements for go.


Use the real browser: clawchrome.com. You can evaluate code on the page, do network capture, screenshot. But it’s the real browser, no fork, no cdp.

For lower RTT you’ll need an instance and proxy close to upstream but a real browser will help with scraping detection.


UUIDv7 as pk is inherently the best option besides just a bigint where you really don’t need a public id.

But a Url62 as a url safe public id from the pk is simple and straightforward to use and comes with few risks of leak issues. Wish postgres had native base62 encoding for url62 now that it has uuidv7 native.


I agree technically but in most use cases the timestamp from uuidv7 is not a security leak. Especially where you’re already sharing that data in some way or another. A default guid is unnecessary if you use uuidv7 I think (in most situations).


It's more for performance that you shouldn't use them as PK's - If you insert a lot you'll get massive fragmentation over time. A sequential Id avoids that and still gives you a unique row.

The Guid is purely for an external system to grab onto something that I can tie back to an actual row in the database but the external system does not need to know anything about the backend other than <guid>.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: