While the idea is very cool, it seems to have most of the drawbacks of traditional regexes (i.e. the little Syntax quirks that eventually one has to learn), with none of the benefits (regex are everywhere including text editors, not just a python thing). It does make them more readable I guess, but I'd like to hear other HNers' opinions.
I think for beginners a better approach would be using something like RegexBuddy (https://www.regexbuddy.com not affiliated, just found it super useful when I started writing regexes).
Regex combinators are quite common in functional languages (where it works quite better than in python). One particular advantage compared to Regex syntax is that you stay inside the language and only use function calls, so you get the documentation, typing and checking property of the language. I said a bit more about that here: https://news.ycombinator.com/item?id=12385840
Mypy is pretty painful to use, though. The benefits of typing aside, patching it into Python has become too-little too-late, and everything feels like a kludge.
You can get a lot of advantages of Mypy by combining it with a whitelist/blacklist (you can do this through a config file). Lots of missing stuff but with strict optionality alone you're gonna catch a lot of real world bugs
it's painful and in hindsight too late but i disagree that it's too little. i've been using it for about a year now and the amount of bugs it finds results in a (very) positive ROI; at $WORK new projects are now started with full mypy awareness from day 1 and old ones are being slowly but surely retrofitted.
I agree with you. At some point, you are bound to get into situations where writing the regex is shorter and easier to do the normal way. Some the examples are getting a little long, and I'm not entirely sure if they are more clear.
It can be ported to py3 with no difficult though it is meant to be used as a tool to construct regex, it shouldnt be put in your programs for a matter of performance at all(mainly those that are critical). About some examples being long, i believe verbosity pays off in understanding complex systems, so yea, in some situations it wouldnt benefit using crocs unless you dont know regex and you dont want to spend some hours to get proficient in it.
I have a similar lib in JS that uses either regexps or plain strings for the leafs, and combinators to implement the various regexp operators.
It is useful for maintainability. The regexp syntax was designed for a write-only scenario (the command line), but complex regexps quickly become unwieldy.
The regular grammar is all about composition, and the regexp syntax (in JS) doesn't allow one to store sub-expressions in variables for reuse and readability.
So this allows one to create regexps piece by piece, to write tests for the sub-expressions, etc...
I really like the way Common Lisp's ppcre library works: the functions all accept either standard strings or a s-expression version of the regular expression and then, using compiler macros, all invocations with statically-determinable regular expressions get compiled to some internal representation at compile-time rather than generating that representation at run-time.
The example is longer than the regex itself however it is simpler to explain to a beginner what crocs's example does than explaning the regex itself. Based on that assumption using crocs's syntax should favour reasoning somehow since it is simpler to understand than obscure regex's syntax.
I think for beginners a better approach would be using something like RegexBuddy (https://www.regexbuddy.com not affiliated, just found it super useful when I started writing regexes).