Messages are reordered with detached sending

#225 · closed · 15 comments

View on GitHub ↗

Quotique

Hello. I'm trying to migrate my project to xtra 0.6.0 (from 0.5.2) and looking for replacement for `do_send_async` method (without awaiting for response). As I understood it must be something like `.send(smth).detach().await`. But now I have a problem with an order of received messages. It is not same as they are sent. Here is example with version 0.5.2 (works): ```rust use async_trait::async_trait; use xtra::prelude::*; use xtra::spawn::Tokio; #[derive(Default)] struct Printer { last_printed: usize, } struct Print(usize); impl Actor for Printer {} impl Message for Print { type Result = (); } #[async_trait] impl Handler<Print> for Printer { async fn handle(&mut self, print: Print, _ctx: &mut Context<Self>) { assert!(self.last_printed < print.0); println!("Printing {}", print.0); self.last_printed = print.0; } } #[tokio::main] async fn main() { let addr = Printer::default().create(None).spawn(&mut Tokio::Global); for i in 1..1_000_000_000 { let _ = addr .do_send_async(Print(i)) .await .expect("Printer should not be dropped"); } } ``` Same with version 0.6.0 (fails): ```rust use xtra::prelude::*; #[derive(Default, xtra::Actor)] struct Printer { last_printed: usize, } struct Print(usize); impl Handler<Print> for Printer { type Return = (); async fn handle(&mut self, print: Print, _ctx: &mut Context<Self>) { assert!(self.last_printed < print.0); println!("Printing {}", print.0); self.last_printed = print.0; } } #[tokio::main] async fn main() { let addr = xtra::spawn_tokio(Printer::default(), Mailbox::unbounded()); for i in 1..1_000_000_000 { let _ = addr .send(Print(i)) .detach() .await .expect("Printer should not be dropped"); } } ``` Do you have any tips for achieving similar behavior?

Comments

thomaseizinger

See #94 for a discussion on this. The tl;dr is that this is by design. Detaching from the response _literally_ means you give up control over scheduling. It is like spawning tasks or threads. If you want things in sequence, wait for the response and then queue the next message. If you don't want to block the current task on that, move the entire loop into a new task: ```rust tokio::spawn(async move { for ... { address.send().await; } }) ```

Quotique

Unfortunately, none of the options suits me. I made a small patch for my problem, maybe it will be useful for someone else: https://github.com/Quotique/xtra/commit/74f20c2edc8d38b8238e249e0a5a0f03156ee63d

thomaseizinger

Can you explain why that doesn't work for your usecase? We thought about this a lot so I am curious to learn where the current model seems to fall short.

Quotique

I use a router actor for load balancing between worker actors. In my case, messages can be grouped, for example, by user id. This way I can have a separate actor to handle individual user requests and a lightweight balancer. Messages from different users are processed independently. But messages from the same user must be processed in the same order. For example, the user can cancel a previous request. So 1. Router by design can't wait for a response from the worker. 2. Reverse reordering on worker side is too complicated. 3. Spawning a task on each message leads to huge overhead on router side. This may become a bottleneck. 4. If I spawn one task to wait for responses, I need a sending mechanism for futures and so on. Maybe I missed something, but it seemed easiest to me to make the patch.

thomaseizinger

> 3\. Spawning a task on each message leads to huge overhead on router side. This may become a bottleneck. I'd like to see benchmarks for this assumption. A task in tokio only incurs a single allocation. You should be able to spawn several thousands a second without problems. You also wouldn't spawn a task for each message, that is redundant. You only need to spawn a task for a group of messages that cares about ordering. Unless you are doing absolutely trivial stuff, the overhead of spawning a single task is neglible in that context. > 1. Router by design can't wait for a response from the worker. Maybe your router shouldn't be an actor? Xtra supports work-stealing by default (it is an mpmc channel). So you can spawn multiple actors that read from the same mailbox (by cloning it).

Quotique

> Maybe your router shouldn't be an actor? Xtra supports work-stealing by default (it is an mpmc channel). So you can spawn multiple actors that read from the same mailbox (by cloning it). Yeah, I know about this option, but my actors are statefull, I need a guarantee that all user messages will be proceed by same actor. Work-stealing delivers message to a first free actor, right? So, my router looks like this: ```rust impl Handler<MyCoolMessage> for Router { type Return = (); async fn handle(&mut self, msg: MyCoolMessage, ctx: &mut Context<Self>) { if let Some(worker) = self.routes.get(&msg.user_id) { let _ = worker.send(msg).await; } else { // Error handling stuff } } } ``` > I'd like to see benchmarks for this assumption. A task in tokio only incurs a single allocation. You should be able to spawn several thousands a second without problems. You're probably right. But tokio scheduler is a black box for me. I'm not sure that everything will fine if I spawn a million of tasks per second. I'll benchmark different approaches later. But now I need to focus on other tasks.

thomaseizinger

> > Maybe your router shouldn't be an actor? Xtra supports work-stealing by default (it is an mpmc channel). So you can spawn multiple actors that read from the same mailbox (by cloning it). > > Yeah, I know about this option, but my actors are statefull, I need a guarantee that all user messages will be proceed by same actor. I can only speak from an abstract point of view but if you have multiple messages that need to arrive in the same order at the same actor, perhaps the messages are too fine granular and whatever you are delegating to your actor should just be one message? Guarding state is where the actor model shines. Actors and their messages are like (micro) services and their APIs. They should always be in a consistent state from the perspective of their callers. They can't control how another service uses their API so they should be able to handle every message in every possible state. If some messages only make sense as a group, I'd consider merging the messages to instead model whatever the group represents as an operation.

DmitryBochkarev

Hi there! I know this issue has already been discussed, but I'd like to share some feedback. First, I want to thank you all for creating such a great library! I've used it in several projects, and it works wonderfully every time. The minimal API and typed messages are fantastic features. For my use cases, I typically use `send` without detach. However, I've noticed that this can effectively limit system throughput based on the target actors' ability to process incoming messages during peak loads. While introducing "casts"(`send(..).detach()`) can reduce response time, this introduces the behavior described in the original issue, which could break systems that assume message ordering. This assumption seems reasonable since in Erlang, multiple messages cast from one process arrive in the same order(relative to each other) in the target's mailbox. I have a feeling that if we could somehow maintain message ordering for "cast" messages, it would greatly benefit the xtra library. Currently, I'm solving this problem by introducing a separate channel(custom struct around Address) that accepts messages in a non-blocking manner. From a user perspective, I would prefer to be able to change the behavior by passing a different MessageBox type. But this would be the ideal case. Thanks again for your excellent work on this library, and I'd be happy to discuss this further(if possible) or help with any implementation ideas if needed!

Restioson

> I have a feeling that if we could somehow maintain message ordering for "cast" messages, it would greatly benefit the xtra library. We did have this but it necessitated having a split between prioritised mailbox and regular mailbox (because of how binary heaps work), where you had extra code for the unprioritised+ordered mailbox and the prioritised+unordered mailbox. You _also_ had some confusion in that one was ordered and one was not (cf. binary heaps comment). I think that it was not seen to be super beneficial when there are other ways to make sure messages arrive in order, compared to the additional API surface and implementation (+ memory weight + complexity) There are three basic patterns that can be used: ## Option 1 - send + await ```rs // pauses future during handling addr.send(Message1).await; addr.send(Message2).await; addr.send(Message3).await; ``` ## Option 2 ```rs // now is a 'cast' but still in order tokio::spawn(async { addr.send(Message1).await; addr.send(Message2).await; addr.send(Message3).await; }); ``` ## Option 3 Send messages prioritised in descending order (sorry I ran out of time to do a code example for this) ------ Are there any reasons why these patterns are insufficient? As I understand it the main downside to reordering is not technical, it is _simply_ that it is a bit confusing (I have sympathy for this, but it was _also_ confusing to have in-ordered sending for non-prioritised but out-of-order for prioritised). But I may not understand the use case

DmitryBochkarev

Thank you for such a fast response! Option 2 could work for some cases. However, in my code I have a strict rule that there should not be any dangling processes - everything should be attached to something. So in my case it would be: ```rust tokio::spawn(xtra::scoped(&addr.clone(), async move { for i in 1..10 { let _ = addr.send(Print(i)).await; } })); ``` With this approach, there are limitations with send+move, which are generally manageable. However, if a process wants to produce batches from 2 different places, this would require some synchronization mechanisms. Here's option 3 (as I see it): ```rust use std::sync::atomic; use xtra::prelude::*; #[derive(Default, xtra::Actor)] struct Printer { last_printed: usize, } struct Print(usize); impl Handler<Print> for Printer { type Return = (); async fn handle(&mut self, print: Print, _ctx: &mut Context<Self>) { assert!(self.last_printed < print.0); println!("Printing {}", print.0); self.last_printed = print.0; } } struct CastAddr<A, Rc: xtra::refcount::RefCounter> { address: Address<A, Rc>, priority: atomic::AtomicU32, } impl<A, Rc> CastAddr<A, Rc> where A: Actor, Rc: xtra::refcount::RefCounter, { fn new(address: Address<A, Rc>) -> Self { Self { address, priority: atomic::AtomicU32::new(u32::MAX - 1), // decrease by one to keep system messages priority as it was } } async fn cast<M>(&self, message: M) where A: Handler<M>, M: Send + 'static, { let priority = self.priority.fetch_sub(1, atomic::Ordering::Relaxed); let _ = self.address.send(message).priority(priority).detach().await; } } #[tokio::main] async fn main() { let addr = xtra::spawn_tokio(Printer::default(), Mailbox::unbounded()); let cast_addr = CastAddr::new(addr); for i in 1..10 { cast_addr.cast(Print(i)).await; } } ``` This is almost what I want. The only issue is that regular sends have low priority by default (or some hardcoded value). Which approach to choose should be determined by business requirements, I think. I believe it would be reasonable for regular sends to use maximum priority, while casts have lower priority.

DmitryBochkarev

Just in case here is code that i currently used(this is for "strong" address, for weak i have the same) ```rust use xtra::prelude::*; #[derive(Default, xtra::Actor)] struct Printer { last_printed: usize, } struct Print(usize); impl Handler<Print> for Printer { type Return = (); async fn handle(&mut self, print: Print, _ctx: &mut Context<Self>) { assert!(self.last_printed < print.0); println!("Printing {}", print.0); self.last_printed = print.0; tokio::time::sleep(std::time::Duration::from_millis(100)).await; } } struct AddressCastChannel<A, M> { addr: Address<A>, tx: tokio::sync::mpsc::UnboundedSender<M>, } impl<A, M> AddressCastChannel<A, M> { fn new(addr: Address<A>) -> Self where A: Handler<M>, M: Send + 'static, { let (tx, mut rx) = tokio::sync::mpsc::unbounded_channel(); let scope_addr = addr.clone(); tokio::spawn(xtra::scoped(&addr.clone(), async move { while let Some(msg) = rx.recv().await { let _ = scope_addr.send(msg).await; } println!("Channel closed"); })); Self { addr, tx } } fn cast(&self, msg: M) { self.tx.send(msg).expect("Channel closed"); } } #[tokio::main] async fn main() { let addr = xtra::spawn_tokio(Printer::default(), Mailbox::unbounded()); { let cast_addr = AddressCastChannel::new(addr); for i in 1..10 { cast_addr.cast(Print(i)); } } println!("All messages sent"); tokio::time::sleep(std::time::Duration::from_secs(2)).await; } ``` This version have limitation - it can cast only one type of messages

thomaseizinger

Unbounded queues break backpressure, potentially causing your system to crash on OOM. If you need to handle bigger bursts of messages, you need to increase your buffer size (mailbox capacity) to sustain throughput. If you are not familiar with the Bandwidth-Delay-Product, I can recommend reading up to it as that is what is effectively at play here, even if we don't have a networked system.

Restioson

> However, if a process wants to produce batches from 2 different places, this would require some synchronization mechanisms. Sorry, can you explain the end usecase a bit more please?

Restioson

@thomaseizinger by the way, the docs for send state > The message will, by default, have a priority of 0 and be sent into the ordered queue do you know if it is out of date or do we need to cut a new release?

thomaseizinger

Not off the top of my head but I think they are probably outdated.