Hacker Newsnew | past | comments | ask | show | jobs | submit | mfateev's commentslogin

Disclaimer: I'm a co-founder of Temporal.

Temporal has 3 types of activities:

* local: executed in the same process as the orchestrator code. Many local activities can be executed locally before their results are sent to a backend server in a single RPC call.

* task queue based: executed by a pool of worker processes that poll from the queue. This is the most flexible model as it supports flow control, priorities, fair queueing out of the box.

* eager dispatch: task queue based but executed locally if possible as a performance optimization.

Temporal also supports stand alone activities that are invoked without a workflow and dispatched through a task queue.

Restate only supports local activities (using Temporal terminology).

I wouldn't call it a "more flexible programming model".

Restate made several decisions I consider questionable for the system's availability and stability, like pushing work to handlers instead of dispatching it through a queue. Any design decision has tradeoffs. It would be nice if you mentioned these trade-offs in your posts instead of making claims that sound like pure marketing.


Sorry, but this is incorrect (Restate founder here)

(1) You can model the equivalent of local activities and activities that run on other workers in Restate.

A local activity is a step in the workflow function. An activity supposed to run on a different worker is a function called by the workflow function. Since these calls are just Restate events, exactly-one, suspendable, this gives you a full-fledged workflow/remote-activity pattern. Including concurrency, separate retry policies, etc.

(2) Restate steps commit individually, unlike local activities.

Imagine a two-step workflow, where you want one step durable before starting the second. Account withdrawal before deposit. Restate steps allow you to do that, each step is durable committed before the next step runs.

Per Temporal's own docs, Temporal Local Activity results become durable only when the enclosing Workflow Task completes. That's different than Restate, which can durably commit every individual ctx.run before proceeding to the next step.

Making actual durable commits fast, so you can have sequences of fast durable steps building on each other is super valuable. If an agent can commit the guardrail evaluation in low milliseconds before it starts the tool call, that's great, do it. If committing this involves dispatching another workflow or activity task, the consideration is harder.

When we see someone migrate a Temporal workflow, they often end up using many more durable steps in Restate than they used activities before.

(3) Why do we consider it flexible?

(a) Virtual Objects: Keep state around across the workflow, without doing tricks like "keep the workflow running, signal only, continue_as_new" after a while. Virtual Objects are a natural way to model concurrent stateful entities.

(b) non-workflow communication patterns: We have seen users build lot's of different patterns. It can get as crazy as graphs of functions/objects sending each other durable messages. All end-to-end idempotent (or exactly once, for Virtual Object state). That is outside the hierarchical workflow/subworkflow/activity abstraction.

(4) availability and stability

I don't know where the perception with stability comes from, Restate pushes some pretty high volumes for customers, like 100k+ actions/sec.

The push model in Restate is internally dispatched through queues as well. The application just don't see it as task queues. Limits are implemented through virtual-queues in Restate 1.7.

-----

Of course the systems make different trade-offs. The Restate design (vqueues, push model) is took us longer to build than a task queue model would have, because it puts more work onto the dispatcher that is otherwise handled just implicitly by the worker pools.

But once it is there, it is sooo nice, in how easy it integrates into infra, and how it can handle flow-control with high hierarchical limits in ways I genuinely haven't seen in any other system achieve (see https://restate.dev/blog/announcing-restate-1-7 )


Might not be par for the course for HN, but gotta drop a link with a light joke, with the background of the money grubbing guy (Temporal founder here) sucking up venture capital like vacuum cleaner who tries to argue he’s intellectually superior in a brag thread about his Company over a competitor (with significant OSS accomplishments) who is attempting to make the world a better place with a freely available primitive that isn’t just available to dev teams at mega teams at Netflix and OpenAI etc.

Stay humble / nice ego dude.

Necessary reference: https://share.google/7uReiRkd6mY3RcCQQ


Hmmm, that feels not very nice, tbh. I have big respect for what Maxim and the team built, and both systems are essentially open and free (publish code and are free to use).


Do you realize that Temporal is fully open source under the MIT License, while the competitor is under BSL?

I certainly don't think I'm superior to them in any way and have great respect for the very capable team they are. We know each other personally. I even presented at the Flink conference while at Uber.

My message is that I'd prefer a more technical discussion of the merits of our products rather than simplified marketing attacks.


Fun fact: The first-ever Durable Execution POC at AWS Simple Workflow used snapshots. The workflow was implemented as a fully asynchronous Java application, and a snapshot was a dump of the whole object graph using reflection. Performance was abysmal because a snapshot had to be generated after every state transition.

So we switched to replay. The biggest benefit of replay is that it lets you implement Durable Execution in any language as a library without a complex runtime. It also supports code changes while workflows are in flight (Temporal calls this patching). Making snapshots of arbitrary code state backward-compatible with code changes isn't practical.

I personally think that, in the long term, Durable Execution will use a runtime that supports both snapshotting and determinism. That way, snapshots can be taken infrequently, and replay can bring workflow code to the latest state. Similarly to a database recovering from a WAL. WASM is the most promising technology to achieve this.


I've built a framework that provides durable threads using serializable continuations in Java (with a modified JVM) and all these issues are easily solvable. Sadly it's not public :( The framework actually does have solutions for those, and all the issues raised in the article, and does support hot patching too. Plus it has a nice waitUntil() API that lets you sleep until an arbitrary combination of events including database query changes.

Performance versus log replay depends a lot on what you're doing, there are plenty of cases where snapshots are faster. Consider anything where you download a lot of data and filter it. But I argue the programming model of log replay is so terrible, and creates so many new classes of subtle bugs, that it's worth paying almost any performance price to get away from it. Especially for durable workflows correctness matters more than performance and it's much easier to achieve with continuation.

In the end it wasn't necessary (because workflows aren't expected to be super fast) but if needed I could have optimized snapshotting further with some more JVM changes. The JVM I was working with is written in Java so is easy to modify.

For hot patching there are a few tricks that help.

1. Only store live data. If a variable points to a large object graph before a checkpoint but isn't used afterwards, don't snapshot it.

2. Make it easy to switch the version of a running continuation only at known-safe checkpoints. I identify checkpoints with a (stack trace, counter) pair.

3. Mostly people want hotpatching only at specific points in their program, typically at the top of an infinite loop that's waiting for something. Design the scheduler so you can expose an API that offers "wait until something happens that I'm interested in, or I change version and then hot swap me", with a test framework that actually drives continuations through those sorts of hotswaps. If you get the API right then you (framework author) control what's live on the stack at that moment and the developer just has to think about the core state of their main root object, which they'd need to think about anyway and is where the important stuff is.

This is better than hotswap/patching with log replay, which is extremely risky - you can't change code at a given point even if you know all your workflows are beyond that point, so it's a leaky abstraction. And you just can't upgrade infinite loops at all, which makes hotswap a lot less useful to begin with.


That sounds like an excellent use for narrow boring company tunnels.


Pretty sure they haven't yet achieved their cost/speed targets, else they would have easily won the contract for this tunnel.


I don't believe in visual programming ever replacing code for general programming. However, many examples exist when visual programs work for narrow domain-specific applications. This works because such applications allow exposing only high-level domain-specific abstractions. This reduces complexity enough to be a good fit for visual representation.


Nah, I wouldn't be too sure. Each generation introduces their own new abstractions on top of whatever existed before. Any code that was written before their era is outdated legacy unreadable stuff, and it can only be fixed by writing new shiny modern blazing fast rocketemoji stuff. It doesn't matter that existing code already works, it must be changed anyway, and we all know that change==progress.

Maybe it won't happen during the next decades. Most likely it won't happen until all currently-alive maintainers of Linux "retire", and most programmers alive are from generations that grew up being pressured to use either Rust or React.

So I can easily imagine a future where there's two main camps: low level code in Rust (where "low level" now means "anything up to, and including, the web browser"); and high level code in some graphical thingy that compiles down to WASM.

And the path I see where we'll reach that point, is by the Rust community continuing to do their thing, and by SaaS companies continuing to do their thing (VM->Docker->"Serverless"[1]->"Codeless"[2]->"Textless"[3]).

With all that, I think it's possible that visual programming might become the dominant way of doing general programming. Not the only way (after all COBOL is still a thing today, so it would be like that), but I can see how what's "normal" is shifted up one level of abstraction higher, so that "code" in that future is seen the same way as "assembly" today; or "assembly" in that future is seen the same way as "punch cards" today.

Those old timers think their Rust language is good enough with their memory safety and stuff, and prefer to ignore decades of progress on programming. Separating the code in text files? Yeah no wonder they keep having incidents like the DeseCRATE or Cargottem supply chain attacks (years 2038 and 2060 respectively), if they install dependencies willy-nilly without even looking at the code they're bringing in (nobody wants to look at dependencies with that primitive tooling). If they used modern tooling instead, they would be able to easily tell at a glance which "nodes"[4] are trying to do suspicious stuff just by their location or relationships with other nodes; or even restrict network or filesystem access to a region of the canvas. They say "well, just use capabilities", but then the code looks like an unreadable mess and way too error prone, when modern tooling lets you just draw a square on a canvas and "any node[4] inside the square can read filesystem, everything outside it can't", without needing to modify any of the nodes[4] themselves. It eliminates whole types of vulnerabilities and whole types of human errors.

(/s, mostly)

[1]: Still needs servers, but you don't worry about it because servers are too low level and you only want to focus on the important stuff (just code).

[2]: Still needs code, but you don't worry about it because code is too low level and you only want to focus on the important stuff (just English).

[3]: Still needs text, but you don't worry about it because text is too low level and you only want to focus on the important stuff (just visual concepts).

[4]: Module, function, macro, etc.


Do you know about the Temporal startup program? It gives enough credits to offset support fees for 2 years. https://temporal.io/startup


I know its gonna sound entitled. But even though we are a small company we still process a lot of events from third parties. Temporal cloud pricing is based on number of actions, 2400 bucks would only cover some months in our case.


If you are expecting to still be small after 2 years that just delays the expense until you are locked in?


temporal.io just released .NET SDK. The observability and scalability of the platform is really good.

Disclaimer: I'm one of the founders of the project.


Nice to see you dropping in Maxim!

For the GP poster - I agree with Maxim here. We've been evaluating workflow orchestration and durable function systems for a while and finally whittled down to where we think we're going to pull the trigger on either Azure Durable Functions or Temporal. Temporal is really nice - the fact that you are "just writing code" is such a huge bonus over some other stuff like AWS Step Functions, Cadence, and Conductor.

As an aside, the engineering/sales engineering team over there seems top notch.


What I meant specifically is that the current state of a workflow is stored in a format that’s opaque to any component other than the workflow itself.

E.g. if I have a “shopping cart checkout” workflow and the user is not making progress, how can I can I tell which step of the workflow the user is stuck at?


Every step of the workflow is durably recorded. So you have the full information about the exact state of each workflow. To troubleshoot, you can even download the event history and replay workflow in a debugger as many times as needed.

The ease of troubleshooting is one of the frequently cited benefits of the approach.

Check the UI screenshot at https://www.temporal.io/how-it-works.


The function's event data and current state is all stored in table storage, so you could query that - I'd expect you'd need to query an event-store-based solution in a similar way?


Check out temporal.io. It has support for schedules as well.


Check out temporal.io that fully abstract this. Disclaimer, I'm one of the founders.


hah, I was thinking about temporal as I was writing this. I have played with temporal pretty extensively.


State machines are useful when the same input/event requires different handling based on the current state. There are not that many applications when this is true. Most of the time only two handlers in each state are needed, success and failure, which are much better modeled through a normal code than an explicit state machine.

At the framework level they might be pretty useful, but they rarely appear at the first version, but as a result of refactoring.


Yes, Temporal workflows are as dynamic as needed.

The other useful pattern is always running workflows that can be used to model lifecycle of various entities. For example you can have an always running workflow per customer which would manage its service subscription and other customer related features.


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: