AstNodeCoverDecl::hier() is relative to the scope the declaration is
emitted from, as V3EmitCImp builds the reported path as 'vlNamep + hierp',
and that is how V3Coverage uses it. V3FsmDetect instead set it to the
absolute scope name, so the instance path was counted twice.
This was masked whenever the owning module was inlined, as the declaration
then ended up in the top scope. With inlining disabled, an FSM in an
instance reported 'top.t.forced_wide_u.t.forced_wide_u'. In the inlined
case it reported 'top.TOP', leaking the internal top wrapper name into
user visible coverage output, rather than plain 'top'.
Nodes were always cloned under the new scope, with the originals deleted
at the end of the pass. Count total module instantiations up front, then
move the nodes instead of cloning them when scoping the last instance.
Required to avoid a memory regression in a follow up, but also faster
and less peak memory overall.
Output is identical.
Each AstVarRef is visited exactly once (blocks are iterated under
their per-scope clone), and the fixups are independent, so no need
for the ordered set.
Store input edges of fixed arity vertices inline in the vertex class.
This reduces heap allocations, memory fragmentation, and pointer chasing
and speeds up Dfg passes. The extra branch introduced in inputEdgep() is
well predictable and profiling shows branchless alternatives are a loss.
Various cleanups, simplifications and improvements:
- Build the output bottom up, eliminating AstSplitPlaceholder
- Replace the 'ignore step' pruning with removing the vertices and edges
- Treat an NBA written variable as a block input, enabling more splits
- Move instead of clone leaf statements
Overall V3Split is faster, uses less memory, can do more splits, is
simpler algorithmically, and has half the lines of code in V3Split
Most notably this can now be split, which could not be before:
```systemverilog
always @(posedge clk) begin
a <= !a;
if (a) b <= c;
end
```
t_x_rand_mt_stability* pin $random values, but there is one random seed
per C thread, so they depend on which thread runs the block. More
splitting moved that block onto a different thread, so force a single
MTask in those tests to keep them stable.
Some property/assert test changed hit counts due to races, but are now
closer to what might be expected.