**Problem**
If a client cannot reach the server, it deletes the portfile. Then it
starts a second server. The first server is displaced: its socket path
now belongs to the second server.
A displaced server keeps running. A dropIfIdle notification cannot
reach it, because the proc file it registered names the socket that
the second server now owns.
**Solution**
The portfile holds the serverId of whichever server wrote it. Watch
that file, and exit when the id is not this server's. Leave the
portfile to the server that owns it.
---------
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
+cleanFull failed with Not a valid key: Full. The cross command re-parses its argument with the key parser, and clean is an accepting prefix, so the rest was taken as task args. Args now have to be preceded by whitespace.
**Problem**
Currently bad resources are silently dropped.
**Solution**
1. Fix a bad resource generator in the build.
2. Fail the build when bad resources are found.
The overload that takes axisValues and settings passes an empty
scalaVersions into the fold that builds the rows, so a call with
autoScalaLibrary = true returns the matrix unchanged. The assertion
records that, so the fix can move it.
Now the call adds one row from the axes it was given, so a build can name a partial or a full version instead.
* [2.x] feat: testForkedWorker setting
**Problem**
A major bottleneck of the default setting is that concurrentRestriction
allows exactly one forked test in flight at a time,
and partly due to the restriction all test suites from a subproject runs
within the forked process.
**Solution**
This adds a new setting called testForkedWorker, which lets the build users
tune the number of concurrent worker process in flight.
I've set the default to CPU count / 3.
Within a single worker, the default parallelism is reduced to 2.
The default test grouping now splits the test suites into
testForkedWorker count.
* Use Global scope for overall worker count
Subproject scoping can be used for single vs split
* Adds math.max(..., 1)
* Avoid running setup/cleanup on an empty selection
* Fix the tag application order
* Rename key to workerMaxInstances
* Implement TestTopology
* fix: Fixes exclusive test
Note due to the limitation of the tag system, setup-clean/cleanup
may overlap.
* Check that workerMaxInstances >= 1
**Problem**
The accept loop did not catch an exception from onIncomingSocket, so
its thread ended. Nothing closed the server socket, so the path stayed
bound: the server took no further client, and the process kept running.
The socket of the client that failed stayed open too.
**Solution**
Log that client, close its socket, and take the next one. An exception
from accept itself keeps the handling it had, so a socket that is
really broken still ends the loop.
The callback now takes a holder rather than the socket, and calls
AtomicCloseable.release to keep it. The loop closes whatever the
callback left behind.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
**Problem**
Nothing covers what the accept loop does when it cannot serve a client.
The next commit changes it.
**Solution**
Assert what the server does now. An exception from onIncomingSocket
ends the accept loop, and the server serves no client after that.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
* fix: restart the sbt server when -D options change
* fix: don't restart on completions, wait for the socket
* fix: only trust -D options recorded for this build
* fix: confirm the server left before reporting it gone
* Update protocol/src/main/contraband/portfile.contra
Co-authored-by: eugene yokota <[email protected]>
* fix: record the server -D options as a list
* fix: record the -D options without their values
The connection file keeps the name of each option and a salted digest of it
rather than the option itself, and the comparison follows the JVM in taking the
last definition of a name. A client that is not allowed to restart the server
says which options it cannot pick up instead of staying quiet about them.
* fix: compare the -D options given before a command
A trailing -D option used to switch the whole comparison off, which dropped the
options written before it, and a shutdown request that never reached the server
waited the full timeout before saying so.
* fix: leave a server no client started alone
The connection file says whether its recorded options are the whole story, so a
server an editor keeps is warned about instead of shut down, while one the client
started with no options is still replaced. A restart that cannot reach the server
warns and carries on rather than failing the invocation.
* build: filter NetworkClient clinit in MiMa
---------
Co-authored-by: eugene yokota <[email protected]>
CommandExchange asked twice whether an exec came from the closing
channel, and cancelled an exec in two places. Each has a name now, as
does the read of shuttingDown and the Try whose failure is ignored.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
Problem
-------
SemanticdbPlugin.compileIncAndCacheSemanticdbTargetRootTask is a cached task
that reads semanticdbTargetRoot as a bare File, so its absolute path is hashed
into the action cache key:
((true,${OUT}/.../classes,${OUT}/.../classes.sbtdir.zip),/tmp/repro/target/out/.../meta)
Every other component of that key is virtualized. This one names one checkout
on one machine, so the entry can never be served to a second checkout, to a
colleague, or to CI. A monorepo we measured had 118 such keys, none of them
reachable across machines.
Fixes#9709
Solution
--------
Read the target root through semanticdbTargetRootVF, following the convention
sourcesVF and resourcesVF already set: an uncached task holding the virtualized
form of a File-typed key, so that a cached task hashes ${OUT}/.../meta rather
than the path behind it. A cached helper would key on the same bare File and
reintroduce the problem one level up.
The scripted test asserts the property the bug breaks: relocating
rootOutputDirectory must leave every key alone, so the second compile is a pure
cache hit. It pins Compile / semanticdbOptions, because -semanticdb-target
carries an absolute path into scalacOptions as well, which is a separate issue
and would otherwise mask this one.
Generated-by: Claude Opus 5
* [2.x] test: Pin how a server writes its files
**Problem**
Nothing covers how the server writes the portfile and the token file.
The next commit changes both.
**Solution**
Assert what the server does now. The tests read the token file only.
The server rewrites it on every authentication, so a test can read it
while the server writes. The server writes the portfile once, as it
starts.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
* [2.x] fix: Write the portfile and zip by rename
**Problem**
A client reads the portfile as soon as it appears. The server wrote it
with IO.write, in place, so a client could read part of one and fail to
parse it. ActionCache staged and renamed its zip by hand.
**Solution**
The portfile calls IO.writeFileAtomically, and the cache zip calls
IO.copyFile, which stages and renames on its own. The token file needs
its staging file kept to the owner, which neither can do yet, so it
still writes its own way.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
* [2.x] sbt-io: upgrade to v1.13.1
* [2.x] fix: Write the token file by rename
**Problem**
writeTokenfile would first remove the existing token file, and then go
through a non-atomic sequence of steps to write a new one. If a client
attempts to read the file during that process, it is likely to find it
either missing or half-written.
**Solution**
Use newly released IO.writeFileAtomically with ownerOnly flag set.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
---------
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
* [2.x] test: Pin how a client meets an unreadable portfile
**Problem**
Nothing covers what the client does when it cannot read the portfile.
The next commit changes it.
**Solution**
Assert what the client does now. The client reads the portfile once, so
a whole one that arrives a moment later comes too late.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
* [2.x] fix: Retry a read before discarding a server
**Problem**
One failure was enough for the client to delete the portfile and start
a second server, which took the socket from a server that was still
running. The client already tried ten times when a Windows pipe was
busy. It tried once when the portfile would not parse, and once when
the connection was refused.
**Solution**
Give those two failures the same ten attempts as the busy pipe. After
ten the client still deletes the portfile, which a portfile left by a
dead server needs.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
---------
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
Follow-up to #9643. The verification added there re-read every downloaded blob right after writing it, and it sat on top of an older cost: DiskActionCacheStore fully re-hashes CAS files on every findBlobs/syncBlobs/getBlobs call, even though the file name is the digest. On a cache hit that meant hashing a whole classpath worth of blobs just to conclude nothing changed.
- Remote blobs (grpc downloads and inline data) are hashed while the stream is written to a staged temp file, then moved into the CAS, so verification costs no extra I/O and tampered bytes never land under the CAS name. Local puts are copied without hashing and are not stamped, so a mismatched local file is still caught by the lookup hash.
- Verified entries are remembered per process by (digest, size, mtime, fileKey), so repeated lookups are a stat instead of a full hash. Any external change to a CAS file invalidates the stamp and forces a re-hash. This is the same trade-off CacheImplicits already makes for file stamps.
The cache now writes each value-distinct ModuleReport once into a "modules"
table and gives each configuration's details a list of indices into it. That is
~5x smaller and hands the reader instance sharing structurally, with nothing to
probe to rediscover it.
CacheStoreFactory.makeCompressed writes gzip-framed JSON and reads either
framing, sniffing the magic bytes so a cache written before a store was
switched to compression still loads. It defaults to make, so existing
CacheStoreFactory implementations are unaffected.
---------
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
**Problem**
Four fields hold a closeable that a later caller replaces. Each one
reads the field, closes what it finds and writes the new value in its
own way.
**Solution**
AtomicCloseable holds such a value. It closes the value it replaces,
and closes the value that loses a race to fill an empty field.
Two things change for a caller. The client replaced its session
without closing the one it dropped, and now closes it. A close that
throws no longer escapes: each of these sites is discarding the value
it closes, and whatever led there matters more than the close.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
This is essentially applying the same fix as in #8795 but for the Windows sbt.bat script.
We now only set -Dsbt.global.base when explicitly asked for with the --sbt-dir option.
**Problem**
1. On Windows, sbt.bat re-parses its own already-received arguments a second time internally, via call/goto with an unquoted variable splice (call :run !SBT_ARGS!, and similarly for -D/-XX flags, several _SBT_OPTS-sourced flags, and the native-client dispatch path). That extra re-tokenization could let &, |, (, or ) inside an argument escape their quoting and run as separate shell commands.
2. The -D/-XX/-- argument handling spliced the raw argument value into a for /F ... in ("%g%") do ( ... ) construct nested up to four levels deep inside multi-line if blocks. Apparently, that confuses cmd to miscount the parentheses.
**Solution**
1. Call :run without the argument.
2. Avoid for /F
A cache-restored tree is a farm of symlinks into the shared CAS. Writing it is
cheap, but walking it is not: resolving each link reads an inode out of a
directory far too large to stay in the OS metadata cache. Measured against a
57 GB / 952k-blob CAS, a cold stat of a CAS blob costs 165 us against 3.6 us for
a regular file, which makes a cold classpath walk of a restored tree roughly 38x
slower than one of real files. On an 81-module build with exportJars, a no-op
compile drops from 14.75s to 5.96s once the tree is real files.
DiskActionCacheStore now seeds the existing symlinkSupported latch from the
filesystem holding the CAS, and materializes entries with copyFile on APFS.
Files.copy reaches clonefile(2) there, so each entry gets its own inode and walks
at full speed while its data blocks stay shared with the CAS: the restored tree
measured 11 MB against 137 MB for the symlinked one. Copies also cannot corrupt
a blob when a task overwrites its output, which a symlink into the CAS can.
On every other filesystem the behavior is unchanged.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
sbt writes two archives per module per build: the packageBin /
packageInternal jar, and the zip the disk action cache stores an output
directory as. Both deflate one entry at a time on the calling thread.
io 1.13.0 adds IO.jarParallel / IO.zipParallel, which deflate entries
concurrently and produce byte-identical archives. Switch both call sites
over, taking the executor from IO.Implicits.zipContext -- io keeps it
behind that object so the choice is explicit rather than an ambient
ExecutionContext.
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
**Problem**
Channels.newInputStream/newOutputStream share a channel-wide lock,
which would deadlock for a duplex communication.
**Solution**
This implements an alternative DuplexChannels functions that's capable
of duplex communication expected of a "socket".
This also duplicates the Java implementation to the worker app,
so we can use JDK domain socket for forked test communication.
**Problem**
ThisBuild scoped baresettings are treated as a common setting,
which results in duplicate appends etc.
**Solution**
Only treat This-project and ThisProject scoped setting as a common setting.
TestRunner now calls endGroup(name, error) before rethrowing an otherwise escaping LinkageError. This completes the TestReportListener lifecycle while preserving the original error propagation and non-zero task result.
Previously, the outer suite catch handled only NonFatal errors. A suite failing with ExceptionInInitializerError could therefore start a listener group and terminate without a terminal callback, forcing integrations to infer failure from rendered output. The direct TestRunner regression checks the error callback, absence of a normal result callback, and rethrow of the same error.
Co-authored-by: Dmitrii Naumenko <[email protected]>
Co-authored-by: Codex <[email protected]>
reboot restarts sbt with a fresh state, so the extra plugin sbt files registered in BasicKeys.extraMetaSbtFiles were dropped and their plugins disappeared. Prepend an early(addPluginSbtFile=<path>) command for each registered file to the arguments handed to the restarted sbt, so they are re-registered.
---------
Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
**Problem**
Commit c8737b8e4f, "refactor: Change the test type" (PR #8181, Eugene Yokota), made Tests.Output and SuiteResult private[sbt] while changing test tasks to return TestResult. The stated rationale was to type test tasks; it did not describe restricting TestResultLogger customization.
TestResultLogger remains a public, documented extension point, but its run method accepts Tests.Output. External plugins consequently cannot implement it in source. This was noticed while working on TW-102637.
**Solution**
Restore public visibility for Tests.Output, SuiteResult, and its companion. Add a scripted external-build regression test that implements TestResultLogger, uses Output.events as Iterable[SuiteResult], runs test, and observes the logger.
---------
Co-authored-by: Codex <[email protected]>
mkInput folded only a 32-bit murmur hash of the task input into the cache key, so distinct inputs collide at the birthday bound (~31 per 500k realistic inputs), silently resolving a task to the wrong cached output. A new DigestHasher hashes inputs into a full-width sha256 Merkle digest. Changes all cache keys, so the cache is repopulated once. Adds a regression test.
**Problem**
InMemoryCacheStore.CacheStoreImpl.read returns cacheStore.read[T]() on a miss
without putting the value in the cache, so the only thing that ever fills the
cache is write. A task whose stored output is already up to date never writes,
so it re-reads and re-deserialises that output on every invocation for the life of
the server, and the cache can never warm up for it.
update is the costly instance: transitiveUpdate runs it once per project in
the dependency closure, and every call goes through
UpdateReportPersistence.readFrom.
**Solution**
Put the value in the cache after a successful disk read. The keying is unchanged
-- an entry is still (path, lastModified), read re-stats the file on every
call, and write still invalidates before rewriting -- so a populated read is
sound for the same reason a populated write is.
Measured on an 81-module workspace, no-op build, real output files (no
cache-restored symlinks), exportJars=true, alternating the two builds over 8
runs each:
update task 398ms across 82 projects -> 30ms across 29
task total 4797ms -> 4623ms
wall (median) 5.15s -> 4.97s
This makes maximumWeight bind where it could not before: reads now admit
entries, so a build whose working set exceeds sbt.file.cache.size will evict
rather than never cache at all. That is the intended behaviour of a bounded cache.
Generated-by: Claude Opus 5
Fixes#899
Add the task join as an extension method on Seq[Def.Initialize[Task[A]]] in
Def. Member resolution tries extension methods before implicit
conversions, so it is selected ahead of the generic conversion regardless of
whether sbt.Scoped is in scope.