Parallel Processing¶
Freyr's parallelism model is built around chunk-level task dispatch. Understanding how work is distributed and how to choose the right iteration method is key to maximising performance.
The parallelism model¶
graph TB
subgraph System["System::Update(dt)"]
Q["CreateQuery()->EachAsync<Pos, Vel>(fn)"]
end
subgraph QueryExec["Query Execution"]
M["Match archetypes by signature"]
C["For each matching archetype:"]
CHUNKS["For each chunk in archetype:"]
TASK["Enqueue chunk task to ThreadPool"]
end
subgraph Workers["Worker Threads"]
direction LR
subgraph Pool0["Worker 0"]
Q0["Queue 0"] --> W0["Worker 0<br/>steals from Q1,Q2,Q3"]
end
subgraph Pool1["Worker 1"]
Q1["Queue 1"] --> W1["Worker 1<br/>steals from Q0,Q2,Q3"]
end
subgraph Pool2["Worker 2"]
Q2["Queue 2"] --> W2["Worker 2<br/>steals from Q0,Q1,Q3"]
end
subgraph Pool3["Worker 3"]
Q3["Queue 3"] --> W3["Worker 3<br/>steals from Q0,Q1,Q2"]
end
end
Q --> M --> C --> CHUNKS --> TASK
TASK -->|LCG hash| Q0
TASK -->|LCG hash| Q1
TASK -->|LCG hash| Q2
TASK -->|LCG hash| Q3
W0 -.->|steal| Q1
W0 -.->|steal| Q2
W1 -.->|steal| Q0
W2 -.->|steal| Q3
When EachAsync is called:
- Freyr finds all archetypes matching the requested component signature
- For each matching archetype, every chunk becomes an independent task
- Tasks are enqueued to per-worker MPMC queues using LCG-based distribution
- Workers pop tasks from their own queue; idle workers steal from others
ExecuteTasks()orWaitForAllTasks()blocks until all tasks complete
Synchronous vs asynchronous iteration¶
Each — synchronous¶
mRegistry->CreateQuery()->Each<Position, Velocity>(
[dt](fr::Entity e, Position& pos, Velocity& vel) {
pos.x += vel.dx * dt;
});
- Runs on the calling thread
- Guarantees sequential ordered iteration (by entity ID within each chunk)
- Safe for cross-entity reads/writes
- No synchronisation needed
EachAsync — asynchronous¶
mRegistry->CreateQuery()->EachAsync<Position, Velocity>(
[dt](fr::Entity e, Position& pos, Velocity& vel) {
pos.x += vel.dx * dt;
});
mRegistry->ExecuteTasks(); // sync point
- Distributes chunks across all worker threads
- Entities are independent — no cross-entity communication within the callback
- Requires explicit synchronisation via
ExecuteTasks()or the registry's built-in sync points - Best for compute-heavy, embarrassingly parallel workloads
| Method | Blocking | Thread pool | Entity order | Cross-entity reads | Use for |
|---|---|---|---|---|---|
Each | Yes | No | Stable | Safe | AI, interactions, debugging |
EachAsync | No | Yes | Unstable | Unsafe | Physics, movement, particles |
Work stealing¶
Each worker thread has its own MPMC queue. When AddTask is called, the task is pushed to one worker's queue using LCG-based distribution:
void AddTask(auto&& func) {
mTaskCounter->AddTasks(1);
mQueueLcgState = mQueueLcgState * LCG_MULTIPLIER + LCG_INCREMENT;
const auto nextQueue = mQueueLcgState % mWorkerQueues.size();
mWorkerQueues[nextQueue]->push(std::forward<decltype(func)>(func));
}
When a worker's queue is empty, it tries to pop from other workers' queues. This work stealing ensures:
- Good load balance even with uneven task durations
- No single point of contention
- Automatic adaptation to heterogeneous workloads
Chunk-level parallelism¶
Each archetype chunk is the unit of parallel work. One task = one chunk.
System::Update(dt)
└─ Query::EachAsync<Position, Velocity>
├─ Archetype A [Position, Velocity] has 3 chunks
│ ├─ Task: chunk 0 (512 entities)
│ ├─ Task: chunk 1 (512 entities)
│ └─ Task: chunk 2 (512 entities)
└─ Archetype B [Position, Velocity, Health] has 1 chunk
└─ Task: chunk 0 (512 entities)
Task count formula¶
For 1,000,000 entities with chunk capacity 512:
More chunks = finer parallelism but higher scheduling overhead. Fewer chunks = less overhead but coarser load balancing.
Overlapping parallel work¶
To maximise throughput, overlap parallel computation with sequential work:
void Update(float dt) override {
// 1. Start parallel physics integration
mRegistry->CreateQuery()->WithLabel("Integrate")
->EachAsync<Position, Velocity>([dt](fr::Entity e, Position& pos, Velocity& vel) {
pos.x += vel.dx * dt;
pos.y += vel.dy * dt;
});
// 2. Do sequential AI work while physics runs in background
mRegistry->CreateQuery()->WithLabel("AI Think")
->Each<AIState>([dt](fr::Entity e, AIState& ai) {
ai.thinkTimer -= dt;
if (ai.thinkTimer <= 0.f)
ai.nextAction = computeNextAction(ai);
});
// 3. Sync — wait for all parallel tasks
mRegistry->ExecuteTasks();
// Now positions are consistent
}
Timeline diagram¶
gantt
title Overlapping Parallel Work
dateFormat X
axisFormat %s
section Main Thread
Schedule Physics : 0, 1
Sequential AI : 1, 3
Sync : 3, 4
section Worker 1
Process Chunk 0 : 0, 2
Steal Chunk 3 : 2, 4
section Worker 2
Process Chunk 1 : 0, 3
Idle : 3, 4
section Worker 3
Process Chunk 2 : 0, 4 Synchronisation points¶
Freyr has implicit and explicit sync points:
Implicit (inside Registry::Update)¶
PreUpdate phase → WaitForAllTasks() + DestroyEntities()
Update phase → WaitForAllTasks() + DestroyEntities()
PostUpdate phase → WaitForAllTasks() + DestroyEntities()
Explicit (user-controlled)¶
Use explicit sync when you need to interleave parallel and sequential work within a single system.
Avoiding dependencies¶
The biggest impact on parallel performance is avoiding dependencies between tasks:
// BAD: Each entity reads data from another entity
mRegistry->CreateQuery()->EachAsync<Position>([this](fr::Entity e, Position& p) {
// This system reads positions from other entities — RACE CONDITION!
auto otherPos = mRegistry->GetComponent<Position>(otherEntity);
p.x += otherPos.x;
});
// GOOD: Independent per-entity work
mRegistry->CreateQuery()->EachAsync<Position, Velocity>(
[dt](fr::Entity e, Position& p, Velocity& v) {
p.x += v.dx * dt; // only reads/writes own data
});
Golden rules¶
- Don't modify archetype structure during iteration — adding/removing components is deferred to
DestroyEntities() - Avoid reading data written by another task in the same frame — use
ExecuteTasks()to create sync points - Don't call
Registry::Updatefrom within anEachAsynccallback — undefined behaviour - Don't throw exceptions from callbacks — behaviour is undefined in parallel execution