Hacker Newsnew | past | comments | ask | show | jobs | submit | Donald's commentslogin

Have you done any math hacking with sol/astra or fable? It’s more fun than using them for coding. The models are great at the monotony, like constructing a Gröbner-basis, etc. But they’re all still absolutely awful at coming up with new ideas, new proof methods, or new constructive forms. So you spend all your time on coming up with novel hypotheses yourself and handing off the rote work to an agent.

It’s also quite fun to get instant results by finding isomorphisms into unfamiliar areas of mathematics that previously would’ve required some networking in order to build a collaborative relationship.


"A mathematician is a person who can find analogies between theorems; a better mathematician is one who can see analogies between proofs and the best mathematician can notice analogies between theories. One can imagine that the ultimate mathematician is one who can see analogies between analogies."

I wonder how models perform on finding analogies between analogies


Well, I suppose that explains monads. It’s one thing to see an analogy and quite another to make it the basis of an API.

> I wonder how models perform on finding analogies between analogies

Load-bearingly verbose, in my experience.


> To feel human, I write code by hand on the weekends.

"Things won are done; joy’s soul lies in the doing." - Troilus and Cressida


I'm getting the same message doing WebGL shader work.


It (Sol, on high) does seem actually really quite competent at this work though (GPU programming). Much more so than the attempts I made with 5.5 earlier today.


Anthropic has made a recent change that Claude Code will only bill extra usage for the print / non-interactive mode:

> Starting June 15, 2026, Claude Agent SDK and claude -p usage no longer counts toward your Claude plan’s usage limits. Your subscription usage limits stay the same and stay reserved for interactive use of Claude Code, Claude Cowork, and Claude.


What does "no longer counts" mean here? Is it going to be free on automated runs? Is it going to be banned entirely on sub and will work with credits only?


Counted separately, and you get an included amount of tokens equal to the cost of your subscription.

https://support.claude.com/en/articles/15036540-use-the-clau...


Billing is a separate issue from whether automation is explicitly supported.


Not in this context, obviously "Claude" support automation as it's also available since always via API, it's obvious poster was talking about subsidized usage.


Isn't this expected if OpenAI models are going to be listed on AWS GovCloud as a part of the Anthropic / Hegseth fall-out?


I was curious what the commenter's business was, and found this post about HTTP protocol latency: https://jacquesmattheij.com/the-several-million-dollar-bug/



>FREEDRUPALWEBSITEHOSTING.COM

Yeah that's not gonna work nowadays.

>DOWNLOADWEBCAM.COM

Is that like Download More RAM?

>BROWSEHN.COM

Hey, I'm browsing that place right now!

>MUZICBRAINZ.COM

This sounds 100% legit no virus softpedia guaranteed.


Gemini 3 Pro Preview gets 96.8% on the same benchmark? That's impressive


And performs very well on the latest 100 puzzles too, so isn't just learning the data set (unless I guess they routinely index this repo).

I wonder how well AIs would do at bracket city. I tried gemini on it and was underwhelmed. It made a lot of terrible connections and often bled data from one level into the next.


> unless I guess they routinely index this repo

This sounds like exactly the kind of thing any tech company would do when confronted with a competitive benchmark.


I mean, the repo has <200 stars, it's not like it's so mainstream that you'd expect LLM makers to be watching it actively. If they wanted to game it, they could more easily do that in RL with synthetic data anyway.


Belated update on this. Gemini reasoning did much better than quick on bracket city today (an easy puzzle but still). It only failed to solve one clue outright, got another wrong but due to ambiguity in the expression referenced and in a way that still fit the next level down making the final answer fairly cleanly solved. Still clearly has a harder time with it than the connections puzzle.


GPT-5.2 might be Google's best Gemini advertisement yet.


Especially when you see the price


My account is just a week apart from yours. Neocities is an excellent project btw, glad to see your passion for the web shine through it.


Talk to federal contractors about their experience with DOGE. They’re literally cancelling contracts and helping to their friends recapture the money as new federal contracts since the money needs to be spent anyway.

The congressional hearings over this are going to be excellent CSPAN viewing.


As someone in NLP who lived through this experience: there's something uniquely ironic and cruel about building the wave that washes yourself away.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: