The PyTorch 2.9 wheels do work. You can pip install torch --index-url <whatever-it-is> and it just works. You do need to build flash attention from source, which takes an hour or so.
I have H100s to myself, and access to more GPUs than I know what to do with in national clusters.
The Spark is much more fun. And I’m more productive. With two of them, you can debug shallow NCCL/MPI problems before hitting a real cluster. I sincerely love Slurm, but nothing like a personal computer.
Your complaint sounds more like the way that you have to access the HPC (via slurm), not the compute itself. After having now tried slurm myself, I don't understand the love for it at all.
As for debugging, that's where you should be allowed to spin up a small testing cluster on-demand. Why can't you do that with your slurm access?
I think that personal computing is more fun than time-shared computing. :)
It's remarkable what can now be done on a whisper-quiet little box. I hope the Strix Halo's will be just as much fun, and they should be, so long as Flash Attention works.
I really don't understand the difference. Either way, it is just a window into a computer. ¯\_(ツ)_/¯
We rent bare metal on-demand and our whole business is to be able to offer compute that you probably wouldn't be able to host in your house $, as if you own it yourself.
So, we made it so that users can get access into the BMC and modify the box however they want. When they are done, we've automated the reset as well. Fully self-service.
$ These boxes are very expensive, weigh 350lbs, sound like a jet engine and consume ~10kW.
100% - slurm is aimed at job maintenance and resource management on HPC clusters. Thus being a pain in the ass for the kind of fast adhoc iteration and testing that AI/ML requires.
Unless you can submit an interactive slurm job and get exclusive access to an H100 for a few hours of dedicated time. If the cluster is overloaded, it’s hard to get those to run when you’d like, but there are still ways. But you do have to be patient.
But it’s still not quite like exclusive access to resources when you want them. So I can see it from both ways.
I don't have a math or computer science background so the more academic publications are almost always quite unreadable to me. I learn a lot by exploring the source code of existing virtual machines instead, and most are written in systems languages.
Sometimes much smarter people than I randomly decide to write articles that democratize access to very complex areas of programming language development. Examples:
Lots of discussion about choice of programming language in the comments below.
- In principle, it should not matter at all, but there are practical reasons why one PL may be better than another in a particular school or context.
- But, all this "choice of PL" discussion is really a discussion about CS1. A CS degree has at least seven other courses -- assuming 1 CS course per semester -- and in practice many more than that. So, if you're going to ask questions about CS1, the question to ask is, "Does CS1 setup students to succeed in the advanced courses?" Classically, these were courses in compilers, operating systems, networking, and so on. These days, you can add distributed computing, machine learning, etc. (but don't subtract the classics).
I think the average American today, including the average admissions officer, has a negative view of technology. So, an application that is unequivocally optimistic about technology is unlikely to be well received. I think that that is what happened here. We also have no visibility into letters of recommendation, which are likely a big factor.
This is an incredibly poor take - curious as to how you just completely made up that ridiculous claim? Ironic as well that you think letters of recommendation matter for college admissions when they are perfunctory for probably > 95% of them. Maybe you shouldn't espouse your opinions on this.
What a PhD student hopefully gets from an advisor is more targeted advice than the type of boilerplate generic advice in this article and others like it.
An advisor who knows what the student wants to accomplish, and is capable of accomplishing should be able to determine when building a working system is more valuable than pushing yet another paper, when reforming science in a small way is likely succeed, and so on.
The problems are not important, but they illustrate failures that are. For example:
- The paper has an example where the model reasons "I'm frustrated" and then produces an answer that it "knows is wrong". You wouldn't know it if you didn't examine the reasoning tokens.
- There are two examples were R1 often gets stuck "thinking forever"
If these failures happen on these questions, where else can happen? We'll start to find out soon enough.