Mmmm idk. I used an LLM to write assembler in my made up API so I don’t think what you say is true. The value in LLMs is precisely that they are not just regurgitating training data, but rather inferring concepts extracted from trained data. If a programming language used concepts completely disconnected from existing paradigms you’re probably right… but that would also be quite challenging for humans to learn to use, since by intrinsic construction it would also be widely separated from human language.
Esoteric languages like BrainFuck are esoteric and difficult precisely because they go out of their way to eschew conceptual links to existing languages or paradigms.
So if you invented a new type of esoteric language with arbitrary syntax and strange operators (not sure how you’d do that, exactly, iirc all fundamental binary operators are known) it might be impossible to use with an LLM even if the user manual was in context… but aside from that, languages and the underlying concepts are extremely generalizable.
Yeah. I have my own file formats for music, pixel art, and levels in a little game I’m building with the kids. LLMs are really good at understanding these proprietary formats that exist nowhere else in the world except on my old laptop.
> I used an LLM to write assembler in my made up API so I don’t think what you say is true.
I think you are underestimating the amount of data/context that is required to use a battle tested general purpose language, for both humans and LLMs. The ecosystem requires official docs, stack overflow answers, blog posts, tutorials, existing source code, subreddits, issue/PR discussions of undocumented features, obscure mailing list threads with rare insights, books, youtube videos, benchmarks, tests suits... and the ecosystem of libraries for the language that also need their own official docs, stack overflow answers...
You also need the collective audit by the community and assurance that this language has been used in production by countless others.
In the age of intelligence on tap, any new language not specifically designed for human-only use will be born into the world with an attendant plague of documentation. And languages are shockingly generalizable. There are few things in any language that cannot be done in any other. Even those are computationally equivalent to some other set of instructions.
What might be the case though, is that new languages will use more context to think about until they are well represented in the training set.
Esoteric languages like BrainFuck are esoteric and difficult precisely because they go out of their way to eschew conceptual links to existing languages or paradigms.
So if you invented a new type of esoteric language with arbitrary syntax and strange operators (not sure how you’d do that, exactly, iirc all fundamental binary operators are known) it might be impossible to use with an LLM even if the user manual was in context… but aside from that, languages and the underlying concepts are extremely generalizable.