Fix unordered data hazards in multi-threaded scheduling (#8133)

The OrderGraph used during V3Order step deliberately omits some variable
accesses from the dependency graph. E.g.: a read of a variable that is
in the reading block's own hybrid sensitivity list emits no edge, nor
does a read ignored due to a force/release, nor an access to a variable
marked 'ignoreSchedWrite' and friends. For serial mode that is fine, the
logic runs one block at a time. In parallel mode two such blocks can run
concurrently, and if one writes what the other reads, that is a data
race at runtime.

These accesses cannot be recovered from the graph edges. They are now
collected from the AST while the OrderGraph is built, and held by the
OrderLogicVertex performing them.

FixDataHazards is reworked around these access lists stored in
OrderLogicVertex, so it is now aware of all variable accesses the logic
makes, including those not encoded by the dependency graph edges. The
previous heuristic of fixing data hazards by merging same-rank MTasks is
removed. Additional edges are inserted instead to prescribe a fixed
ordering of conflicting MTasks. To insert edges without unduly
increasing the critical path, or introducing cycles, new edges are
added such that they preserve topological ordering, and they are
inserted between vertices sorted by critical path length. See algorithm
details in the code.

Also add a data hazard checker under '--debug-partition', reporting every
unordered accessor pair left in the final MTask graph.

This fixes the race demonstrated by t_sched_hybrid_hazard (#7913),
which is no longer expected to fail.

Under ThreadSanitizer over the vltmt tests: 17 failing before, 3 after,
with no regressions. The 3 remaining are different defects.
This commit is contained in:
Geza Lore
2026-08-18 08:50:50 +02:00
committed by GitHub
parent 96ea587df0
commit d4a18d4dfb
10 changed files with 477 additions and 259 deletions
+72
View File
@@ -261,6 +261,78 @@ void OrderMTaskGraph::mergeMTasks(LogicMTask* recipientp, LogicMTask* donorp) {
VL_DO_DANGLING(donorp->unlinkDelete(this), donorp);
}
void OrderMTaskGraph::removeTransitiveEdges() {
// Removing a transitive edge cannot change any critical path, so none need updating here.
// Only the edge heaps and the dependent sets need maintaining.
for (V3GraphVertex& vtx : vertices()) {
for (V3GraphEdge* const graphEdgep : vtx.outEdges().unlinkable()) {
MTaskEdge* const edgep = static_cast<MTaskEdge*>(graphEdgep);
LogicMTask* const fromp = edgep->fromMTaskp();
LogicMTask* const top = edgep->toMTaskp();
// If the MTasks are also connected by some other path, then this is a transitive edge
if (!pathExists(fromp, top, edgep)) continue;
// Maintain the additional data structures of the OrderMTaskGraph
fromp->removeDependent(top);
fromp->removeRelativeEdge<GraphWay::FORWARD>(edgep);
top->removeRelativeEdge<GraphWay::REVERSE>(edgep);
VL_DO_DANGLING(edgep->unlinkDelete(), edgep);
}
}
// Confirm the above left the maintained state consistent
validate();
}
void OrderMTaskGraph::removeEmptyMTasks() {
// This transform preserves the critical paths as it connects every predecessor of the
// removed MTask to every successor, and the removed MTask itself has zero cost.
for (V3GraphVertex* const vtxp : vertices().unlinkable()) {
LogicMTask* const mtaskp = static_cast<LogicMTask*>(vtxp);
// Keep the entry and exit vertices.
if (mtaskp == m_entryp || mtaskp == m_exitp) continue;
// Keep any MTask that holds logic
bool empty = true;
for (const OrderMoveVertex& mVtx : mtaskp->vertexList()) {
if (mVtx.logicp()) {
empty = false;
break;
}
}
if (!empty) continue;
// The MTask holding no logic should have zero cost
UASSERT_OBJ(!mtaskp->cost(), mtaskp, "MTask holding no logic should have 0 cost");
// Connect each predecessor directly to each successor.
for (V3GraphEdge& inEdge : mtaskp->inEdges()) {
LogicMTask* const fromp = static_cast<MTaskEdge&>(inEdge).fromMTaskp();
for (V3GraphEdge& outEdge : mtaskp->outEdges()) {
LogicMTask* const top = static_cast<MTaskEdge&>(outEdge).toMTaskp();
if (!fromp->hasEdgeTo(top)) addEdge(fromp, top);
}
}
// Remove incoming edges of 'mtaskp'
while (MTaskEdge* const edgep = static_cast<MTaskEdge*>(mtaskp->inEdges().frontp())) {
LogicMTask* const relativep = edgep->fromMTaskp();
relativep->removeDependent(mtaskp);
relativep->removeRelativeEdge<GraphWay::FORWARD>(edgep);
VL_DO_DANGLING(edgep->unlinkDelete(), edgep);
}
// Remove outgoing edges of 'mtaskp'
while (MTaskEdge* const edgep = static_cast<MTaskEdge*>(mtaskp->outEdges().frontp())) {
LogicMTask* const relativep = edgep->toMTaskp();
relativep->removeRelativeEdge<GraphWay::REVERSE>(edgep);
VL_DO_DANGLING(edgep->unlinkDelete(), edgep);
}
// Delete the empty MTask
VL_DO_DANGLING(mtaskp->unlinkDelete(this), mtaskp);
}
// Confirm the above left the maintained state consistent
validate();
}
// Check the critical paths in the given direction, and the critical paths cached in the edge heaps
// in the opposite direction, against those implied by the edges. Note this deliberately iterates
// the edge lists, rather than consulting the edge heaps, so the heaps are validated, not trusted.