Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This doesn't seem significantly different to git, from the perspective of a user. Under the hood, I model a git branch as a ref pointing to a commit object in a DAG. In the driver's seat (git log --patch branch..upstream), I think of it as an ordered collection of diffs.


In darcs or pijul there is no DAG; there is only a collection of patches. The ordering between them is implicit, computed on demand, and can change as a consequence of merges.


It sounds as this is mostly an internal difference. Certainly, the mental model you described can be used on git reasonably well, from the user's point of view, and they won't get steered too far off course with it.


It can be used in git, with a lot of duct tape like git rebase and git cherry-pick and occasional dynamiting through rough git merges.

The darcs/pijul world of repositories as loose ordered sets of patches makes cherry-picking the rule rather than the exception. A branch is mostly just the subset of patches you are interested in at a given moment. A "trunk" is just the superset of all possible patches. You can make interesting and easy usages of things like set intersections [1]: the intersection of the patches in two branches in darcs/pijul can be much more interesting than nearest common parent commit in the git DAG, and especially can be a lot more informative in the cases where things like bug fixes are cherry-picked across branches, which in git is a special bit of tree/commit surgery but in the darcs/pijul world that patch can be often the exact "same" in both branches.

[1] Aside, I love the concept of using intersection branches for consensus-oriented development (what releases to Production are the patches that every developer has pulled into their own working branch), which is a neat form of decentralized development that I think can only really be handled in the darcs/pijul model. (I have an ancient blog post on the idea of such a starfish development workflow.)


There are definitely many implications for users.

One thing you'll notice very quickly with Darcs (and presumably Pijul) is that the system always manages dependencies between patches. If you try to cherry-pick a single patch from a branch, you will get that patch and all the patches it depends on; you don't get the full linear history, you only get a subset.

In other words, you get the intuitive feeling that you're operating not on a log, but on a graph. Pulling one thread necessarily pulls other threads, and the whole graph rearranges itself to accomodate your changes. This has its downsides compared to the strictly-linear snapshot model, but the upside for most users is incredible. You can just commit and merge, and the system handles ordering for you.

Git was a major step down, UX-wise, when we switched from Darcs back in 2008, and it's still less user-friendly today. (It was also a major step up in some ways: Darcs, at the time, had a huge performance edge case where conflicts where sometimes effectively unresolvable because they took too much time to compute.)


I miss Darcs too. One potential downside of making cherry-picks very easy is that abusing them can bite you: two patches might be independent in terms of conflicts but functionally dependent (e.g., code in patch A calls a function introduced by patch B).


It's not quite the same.

Here is a very short video on a project called "Camp" that stalled out, nearly a decade ago. I think it very nicely explains how the user interface differs:

https://www.youtube.com/watch?v=iOGmwA5yBn0

The most important thing is that because there is no DAG, when you say "Darcs, pull this patch for me", like saying `git cherry-pick ABCDEF` -- the dependencies are automatically computed and pulled as well. You can sort of imagine it like if you had a git branch, and you ran 'cherry-pick' on one of the commits to your 'master' branch (because you wanted it). But, rather than pulling that one thing, 'cherry-pick' implicitly traversed the dependent patches and picks them all as well. But because there is no DAG, a dependency doesn't mean "parent commit". It means "the other patches that are mathematically required for this patch to work out". That means cherry-pick always works: you never have to calculate the dependencies yourself. To merge a patch is to implicitly merge all of its dependencies.

I've spent plenty of my time as an OSS maintainer dealing with merging multiple bug fixes from a development branch into stable branches. For example, a bug may already be fixed in HEAD when it's reported, but not STABLE, so you want to pull changes from HEAD into STABLE. Many times this requires multiple, carefully curated sequences of 'git cherry-pick' in order to correctly get the dependencies right. For example, the author may have made a small refactoring, then implemented the bugfix on top of that. Or it requires a complete reformulation or re-commit of a new change that matches the STABLE branch.

In a sense: this never happens with Darcs. If there's a bugfix, I say "Get me that bugfix patch". It always gets every dependent patch that is necessary, and never anything more. Every time. It always just works. Remember: no DAG. You aren't traversing parent commits. You are, in a sense, finding the transitive closure of "patches that cannot commute with this patch" (IIRC). That means: if a given patch does not commute with this patch, i.e. it is dependent, because we must apply them in a certain order, so there is a dependency -- then you also need that patch. And you need to apply that rule to that patch, and every patch it depends on, and so on and so forth (hence 'transitive closure')...

This allows a very powerful form of development, where features and bugfixes can coexist. But they do not necessarily need separate 'branches', so to speak. To merge a feature into a repository implicitly pulls its dependents, and the same with bugfixes. The net effect of this is that Darcs almost always gets merges correct, or it fails to do the merge at all. This kind of means that merges are sound ('kind of' because I don't know about an actual soundness proof, but the intuitive idea roughly is right): if Darcs pulls off the merge, then it's always correct, but it may not be able to always actually do that merge (perhaps not every merge is actually sensible, in the theoretical view of things, or perhaps the merge is sensible but the model doesn't allow it to handle that case).

The "fails to do so" is the tricky part, where Darcs 1 originally went exponential in some cases, though Darcs 2 mitigates this. It looks like Pijul will finally nail this problem dead, although admittedly I haven't looked over the theory.

Side note: Camp was originally envisioned to be the successor to Darcs, or at least the basis for "Darcs 3", using Coq to build formal proofs about the underlying patch theory to show it worked out correctly and avoided the harry bits that plagued Darcs 2. Unfortunately, it never panned out that way (due to time and lack of funding). The project was actually started by Ian Lynagh who worked at my current company before me and was one of the founders.


> If there's a bugfix, I say "Get me that bugfix patch". [..] It always just works.

> Darcs almost always gets merges correct, or it fails to do the merge at all.

One of these is not like the other, which IMO is the problem with "magical" merging systems. Great when they work, f*cking hell nightmare when they don't.

I'd rather have something like git that works in normal usage all the time, and when it fails, is easy to fix. YMMV.


If you try to analyse merge systems mathematically, git's merge system is the "magical" one. It is just a heuristic algorithm, with no solid property you can rely on.

In contrast, Darcs and Pijul's merge are associative, and Pijul's merge is commutative. Even if you don't like maths, this means that they will always behave deterministically. This also means you can use them in scripts, although darcs might sometimes have performance problems (pretty bad ones, actually).

In git, you can get the following: https://tahoe-lafs.org/~zooko/badmerge/simple.html


FWIW, this is traditionally quite possible with Darcs. I sort of misformulated in my original post; it's not like it just gives up and you're at square one. The workflow IIRC was basically the same as Git: it'll throw its hands up and you make a new commit to fix everything. So that isn't really a problem or any different. Note that Darcs 1 did have the exponential merge case on top of this, however, which was pretty unfortunate (and really a byproduct of the design of the change format, among other things).

In all honesty, given years of experience with Git, and fondly using Darcs as my first version control system: I still think merges are absolutely the one thing it beats Git at, hands down. When it works and it does its job, it always is correct. When it doesn't, you can bail it out. Not much different, but the "always is correct" and dependencies-being-implicit is what makes it good. Darcs could have saved me at least dozens of hours of hair pulling when doing STABLE merges I estimate... Git's still good. I wish it could do that, though...

Your note about git is interesting. In fact, Git is, in at least some cases, more magical than other VCSs in the merge department. You might just not be aware of it due to being so familiar. When I say "Darcs always gets the merge correct", I don't just mean it literally finishes with exit code 0, but also that the semantic model is, in some sense, more 'correct' or 'intuitive':

http://r6.ca/blog/20110416T204742Z.html

Darcs (and others) always get this 'merge associativity' case correct, where 'Base+A+B' where (+) is merge is associative (so it doesn't matter how you 'bundle' the changes or whatever). That means you have less edges to worry about. And to be fair, I don't think there's anything inherent about Git where this particular case can't be fixed. It's just a good example of why people are trying projects like Pijul/Darcs at all, so these things can be formalized and understood. The theory of patches is actually rather rich and helps formalize a lot of these notions of what a "merge" really is in an algebraic sense, how patches relate to one another, etc.


It is different. When you commit in darcs, you cherry pick as default.


I don't understand, you will have to use more words. I've never used darcs, and as a git user I don't see how you could cherry-pick by default for commiting. By default, I would say you… create a commit object with the author/date/message/tree/parent metadata recorded in it.


When you type git pull, it has a remote and a ref at that remote, and it attempts to merge (ff notwithstanding) the DAG you have with the DAG the remote has.

When you type darcs pull, darcs lets you pick and choose which patches (subject to the dependency constraints) you want to pull in. Those patches then get applied however darcs wants to apply them (again, subject to dependency constraints, of cousre). Because you do not necessarily need to pull all of them, you are always "cherry picking" by default.

This is different from cherry-picking in git, because when you cherry-pick in git you still have 2+ commits that exist in the context of a DAG; you're just transplanting the contents elsewhere to create a new commit in a different position in the DAG.


I still don't get it: pulls don't merge DAGs, they add objects and move refs (and maybe merge); every time darcs fans say 'patch' (as a collection of related changes) I think, 'oh like a branch?'

This sounds less and less like a tools/implementation thing and more like the default recommended/enforced workflow thing.


To start by adding to your confusion: one of the early problems encountered in darcs<->git interaction actually was that sometimes a darcs patch acts more like a git branch than a git commit...

It might help to compare the "identity" structures of git commits versus darcs/pijul patches. In a pseudo-C, you can see a git commit as something like:

    struct commit {
      string author;
      string description;
      tree_id tree_snapshot;
      commit_id[] parent_commits;
    }
If any of those fields change (are amended), you have a new commit.

This is a directed acyclic graph (DAG) because of that `parent_commits` link from one commit to the immediate previous parents (it can be multiple parents in the case of a merge commit). Git can only just move refs on a pull in the case of a "fast forward" when the remote branch is "simply" ahead of the current branch and all of its new commits "point to" the last commit in the current branch. Every other case it is a merge of the graph (via a merge commit with two or more parent commits).

(While git outputs a diff as the representation of the commit in places like `git show`, a commit doesn't store the diff but instead a link to a snapshot of the tree at the time of the commit.)

For something like darcs/pijul, the identifying information of a patch looks something more like:

    struct patch {
      string author;
      string name;
      change[] changes;
    }
If any of these things are changed (amended) you have a different patch.

This may seem like semantic quibbling in that the patch here actually contains the diffs as a part of its identity rather than a snapshot of a source tree, but that's not actually the important difference.

The important difference is that the context of the patch is no longer a part of its identity: there is no "parent patch" information, and the change structures don't directly refer to previous changes.

The reason that difference matters is because in the darcs/pijul models the context of the patch is more "metadata" about the patch than a direct part of the patch. Patches aren't "nailed" to a graph like a commit is, they "float in a basket" together. Darcs and pijul do the work to figure out which patches need to be in which order in a branch/repository.

This can be a nightmare to someone expecting a strict graph. Darcs and pijul can and will reorder history during a pull. You can see "newer" patches float down under "older" patches in the patch log as the systems work to build a stable sort of patches.

That movement, however, is also where the systems draw the most strength. That movement of the patches can be seen as a continual, rustling "cherry arranging" as the systems work to figure out the minimal set of previous changes that patch needs in order to exist.

If you cherry-pick a commit in git you copy the changes from that commit to the new branch (a new spot in the DAG) into a new commit with its own new identity. Down the line when you go to reintegrate/remerge the branches between the original branch and the cherry picked branch, git doesn't see the same commit/change and its merge can (in my experience, will) see conflicts in the exact same change made in different contexts.

When you cherry pick a darcs/pijul patch, you bring over the same exact patch and the system lets you know any other minimal dependencies that you need and brings them over as well. When you reintegrate/remerge the cherry picked branch, those exact same cherry-picked patches are already "in" the original branch and so don't necessarily need to be remerged/rearranged again.

You can duplicate git workflows on top of darcs/pijul, but it is very hard to duplicate some of the more interesting darcs/pijul workflows on top of git. Among other things, rebase/cherry-picking merge hell is a very real problem in the git ecosystem, whereas darcs/pijul almost seem like crazy smart magic in comparison when it comes to some of the scenarios where you might rebase or cherry-pick.

It might be something that won't entirely make sense until you try experimenting with it yourself: maybe, you might want to take darcs for a spin for a small project or two. I think you can feel a lot of the difference as you use it, especially as you start to push/pull between branches/repositories.

(Anecdotally, my workflows are quite different on darcs versus git, knowing that typically I could fix a bug discovered elsewhere in the code in the middle of a bigger project, without needing to branch, I would often just record that change into a tiny patch on its own right there on the spot, and generally know that if I needed to get just that one patch into another branch I could rely on darcs to cherry pick it for me later.)


  Down the line when you go to reintegrate/remerge the
  branches between the original branch and the cherry picked 
  branch, git doesn't see the same commit/change and its merge
  can (in my experience, will) see conflicts in the exact same
  change made in different contexts.
Finally I see some common ground: by tracking changes separately to commits, users won't see merge conflicts when the commits get moved.

Sadly, git also recognises this on the user's behalf, so likely your experience was due to some other delightful quirk of the git UI.

edit: I'd also recommend not calling trees 'tree snapshots', because that will confuse people familiar with trees. Same for 'cherry-picking a commit', since from your description darcs seems to use 'cherry-picking' to mean 'fetch a set of changes from someone else', which maps to 'fetch a branch' in git-land. 'git cherry-pick' means, 'copy a single change from one local branch to another', so has almost no overlap.


Obviously there are plenty of anecdotes on both sides, but I had to (try to) force a moratorium on git cherry-pick between long running branches at a previous job because merge problems became a huge sink of time. (This was after they'd already had the same problems in TFSVC and seemed adamant to recreate that problem in git.)

You seem to think its "fetch a branch", but I'm trying to tell you that the `darcs pull` experience is a lot more like doing `git fetch && git cherry-pick origin/TIP --interactive` every time than `git pull`, but with a much, much better merge experience than that implies.

I'm not trying to confuse different concepts, I'm trying to show that the hard concept in the git case was the easy concept in the darcs case.


How does the fetch part know which objects to grab? (Answer: because it cared about the DAG first.)


That's begging the question: git-upload-pack takes a ref and a commit ID, and returns a pack of objects to the requester to do as they like. I wouldn't call it a merge, but a set of operational transformation actions to be added to the requester's repo. It may not even examine the DAG at all, since it's possible to save those packs as 'bundles'!

My point is, if you focus on the implementation differences, my understanding won't increase because we don't have the same mental model of how git works under the hood.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: