Arma Reforger Workbench workflow, testing guidelines, and debugging patterns
Quick reference for working with Arma Reforger Workbench. For detailed patterns, see resource files below.
Use this skill when:
Two automated gates β but they cost very different things. tools/compile-check.sh (compiles all EnforceScript, ~5s, headless) is free: run it constantly. tools/run-tests.sh (runs Overthrow's autotests in the real game client, ~15-19s) launches a Reforger window that steals the user's desktop focus, so it is a scarce, deliberate regression gate: orchestrator-only, once after a phase or fix is complete, never during planning, never inside a subagent β see .claude/test-policy.md. Coverage is a spine, not the surface β 30 assertions across four tiers (pure logic, manager init, started campaign, same-session persistence), reachable as one command. Not covered at all: JIP/multiplayer (the most common regression class β it needs two client processes), UI, performance, AI movement, and the save/reload round-trip (written but gated behind the persistence migration). Everything in that second list is still verified by manual Workbench play-testing; be specific about what the user should test and how.
See: testing-guidelines.md for manual test procedures Β· "Writing Autotests" below for automated ones
EnforceScript has specific error patterns. Common issues: missing semicolons, ternary operators, type mismatches, missing strong refs. tools/compile-check.sh surfaces them as file:line: message on stdout; the same errors also show in the Workbench console during interactive sessions.
See: compile-errors.md for common errors and fixes
Use Print() for debug output, check Workbench console logs. No interactive debugger. Add debug prints strategically to trace execution and inspect values.
See: debug-patterns.md for debugging techniques
Prefabs edited in Workbench, layouts in UI editor, configs in text editor. Save often. Workbench can crash. Always test changes in play mode.
See: workbench-tips.md for Workbench best practices
tools/compile-check.sh after code changes; do not ask the user to compile for you (see tools/README.md)tools/run-tests.sh "{6A6E29FF47ECB840}" (Fast) opens a Reforger client that steals focus. Run it once, in the main thread, after a phase or fix is complete and only if the change touches tested areas β never to take a baseline, never from a subagent, never mid-change. Full rules: .claude/test-policy.md Β· contract: tools/README.mdtools/compile-check.sh yourself (~5s warm)file:line: message outputtools/run-tests.sh "{6A6E29FF47ECB840}" once, at the end (Fast, ~16s) β or the All group if you touched campaign, economy or persistence state. Not after every edit in step 2β4: this opens a Reforger client and takes the user's screen. Skip it altogether for docs/layout/prefab/localization-only work, and defer it if the user is play-testing (.claude/test-policy.md)Run the compile check yourself β do not ask the user to compile:
tools/compile-check.sh
Once it exits 0, hand off runtime testing to the user:
Please test in Workbench:
1. Enter play mode
2. Test: [specific steps]
3. Report: [specific things to check]
Read the compile check's stdout β each line is file:line: message. Fix the errors and re-run until exit 0. Do not ask the user for compile errors; the check is the source of truth. (Syntax errors abort at the first failing file, so fixing one may reveal more β just re-run.)
Runtime errors and debug prints still come from the user. Ask:
Please:
1. Check the Workbench console for any error messages
2. Try: [specific test step]
3. Report: What happened vs what should happen
tools/compile-check.sh # compile all EnforceScript headlessly (~5s warm)
tools/compile-check.sh --help # all flags and exit codes
Exit codes: 0 = verified clean Β· 1 = compile errors Β· 2 = indeterminate/tool failure (never means pass or fail) Β· 124 = timeout.
Output: errors on stdout, one per line, gcc-style and repo-relative:
Scripts/Game/Components/OVT_Foo.c:214: error: Broken expression (missing ';'?)
Summary on stderr; full engine log kept at .tmp/compile-check/last.log. Full contract (flags, env vars, limitations): tools/README.md.
Also available: tools/launch-game.sh launches the game client with Overthrow loaded and reports the run's log directory. tools/run-tests.sh is built on it β prefer run-tests.sh and do not drive the launcher directly for test runs.
Before you run anything: this command opens a real Reforger client and takes the user's screen for ~15β20 s. It is the orchestrator's gate, spent once after a phase or fix is complete β not a baseline, not a subagent's job, not a progress check.
.claude/test-policy.mdhas the full rules; if you are not sure you are allowed to run it, you are not.
tools/run-tests.sh "{6A6E29FF47ECB840}" # Fast group β Logic + Init, 18 cases, ~16 s
tools/run-tests.sh "{6A6E2A002F53A581}" # All group β + Campaign + Persistence, 30 cases, ~19 s
tools/run-tests.sh OVT_TEST_LogicSuite # one suite (debugging)
tools/run-tests.sh OVT_TEST_Logic_Town_SupportPercentage_Boundaries # one case, inside its owning suite
tools/run-tests.sh OVT_TEST_MetaSuite # red-path proof β MUST exit 1
tools/run-tests.sh # default target: OVT_TEST_SmokeSuite
tools/run-tests.sh --help # all flags
The two group GUIDs are a stable contract β quote them verbatim (the braces need the quotes) and never read the .conf files. Run Fast after most changes; run All before handing work over or when you touched campaign, economy or persistence state. Suite execution order inside a group is alphabetical by class name, so no case may depend on another.
Exit codes: 0 = all passed (positively verified from junit.xml) Β· 1 = test failures (names on stdout) Β· 2 = indeterminate/tool failure β including a mistyped target, which produces no junit.xml Β· 124 = timeout (default 300 s).
Artifacts land in .tmp/run-tests/: junit.xml, autotest.log, autotest_failed.log, console.log (+ crash.log if the client crashed). Verdicts come from junit.xml only β a green run still prints some SCRIPT (E) lines from gameplay code, and a campaign start emits 62 known VM exceptions, so never judge a run by console error counts.
Save-state precondition. Campaign- and persistence-tier runs assume a fresh save DB:
.scripts/reset_save.sh --profile OverthrowCI # NEVER without --profile
Without --profile OverthrowCI (or an explicit OVERTHROW_SAVE_DIR) these tools target the user's real Workbench campaign save β an rm -rf on hours of play. The guard refuses implausible paths and prints what it resolved, but do not rely on it. Today the reset changes no verdict (nothing writes a save on this branch); it is load-bearing for the acceptance gate below. All three save tools: tools/README.md β Save-state control.
Persistence acceptance gate. OVT_TEST_PersistenceRoundTripSuite is quarantined, in no group, and red on purpose: tools/run-tests.sh OVT_TEST_PersistenceRoundTripSuite exits 1 with Persistence capability absentβ¦. Exit 0 means the core/persistence migration is complete. Do not "fix" it, do not add it to a group, do not weaken it.
Full contract: tools/README.md. Empirical ground truth (verbatim artifact shapes, timings, framework gaps, valid for Reforger 1.7.0.54): docs/features/dev-ops/autotest-foundation/findings.md and docs/features/dev-ops/test-coverage/findings.md.
Pick the tier first. Suites are organised by setup cost, not by subject β the world transition and the campaign start are paid per suite, so coverage grows by adding a case file to an existing tier, never by adding a suite. Choose the cheapest tier that can express the assertion:
| Tier | Suite | Available to a case | Put a case here when it⦠|
|---|---|---|---|
| A Logic | OVT_TEST_LogicSuite |
nothing β no world, no game mode, no manager (engine-enforced: GetWorldFile() is empty, so the game-mode getter is null) |
is pure computation on hand-built objects: record maths, modifier recalculation, a job condition, a skill effect, a level curve |
| B Init | OVT_TEST_InitSuite |
world + managers, campaign not started | asserts something true at world load: a manager resolves, towns are populated, controllers are registered, a config-driven price seam round-trips |
| C Campaign | OVT_TEST_CampaignSuite |
started campaign | needs campaign-start products: activated towns, stocked shops, income calculators, faction state |
| D Persistence | OVT_TEST_PersistenceSuite |
started campaign + the one save seam | writes state through a manager's public mutator and reads it back through its public accessor |
| D' RoundTrip | OVT_TEST_PersistenceRoundTripSuite |
β | nothing new goes here. Quarantined migration gate; red by design |
Prefer Tier A wherever the logic allows β those cases cost nothing (14 of them run in 6-9 s) and cannot flake.
Where files go
Scripts/Game/Tests/
βββ TestFramework/ glue β modded SCR_AutotestHelper (world + addon list), OVT_TEST_SuiteBase
βββ TestSuites/<Tier>/OVT_TEST_<Tier>_<Subject>.c a case file in an existing tier directory
Naming: suites OVT_TEST_<Tier>Suite; cases OVT_TEST_<Tier>_<Subject>_<ExpectedBehaviour>. The class name is the CLI argument, so keep it greppable. Case execution order inside a suite is alphabetical by class name β no case may depend on another having run, or leave state a later one needs.
Skeleton (a new case in an existing tier is the normal job; the suite line is shown for context)
//! Suites inherit OVT_TEST_SuiteBase β never SCR_AutotestSuiteBase directly.
//! [BaseContainerProps()] is MANDATORY or group configs silently instantiate nothing.
//! The world comes from the modded helper β only the pure-logic tier overrides GetWorldFile().
[BaseContainerProps()]
class OVT_TEST_CampaignSuite : OVT_TEST_SuiteBase
{
//! Opt in to a started campaign. Default is false; the guarded start
//! sequence itself lives in OVT_TEST_SuiteBase and is a TEST concern only.
override bool RequiresStartedCampaign() { return true; }
}
[Test(suite: OVT_TEST_CampaignSuite, timeoutS: 30)]
class OVT_TEST_Campaign_Economy_PricesAreInitialised : SCR_AutotestCaseBase
{
protected int m_iTicks;
[Step(EStage.Setup)]
void Setup()
{
Print("setting up"); // shadowed Print -> routes to autotest.log
}
[Step(EStage.Main)]
bool Main() // bool step: re-run EVERY tick until it returns true
{
m_iTicks++;
if (m_iTicks < 5) return false; // wait for managers to settle
OVT_EconomyManagerComponent economy = OVT_Global.GetEconomy();
if (!AssertTrue(economy != null, "economy manager missing")) return true;
SetResultSuccess();
return true;
}
[Step(EStage.TearDown)]
void TearDown() {}
}
Rules that bite
void steps run once. bool steps are re-run every tick until they return true. Returning false forever = a hang bounded only by the step timeout.[Test(timeoutS: N)] / [TestStep(stage, timeoutS: N)]. [Step(EStage.X)] is the short alias of [TestStep].maxAttempts: on [Test] is banned. It exists in the framework and appears nowhere in Overthrow's test tree. A test that needs retries is a bug in the test, a race, or a real defect β fix the cause. Using it requires recorded evidence of genuine engine-timing non-determinism in findings.md.docs/features/dev-ops/test-coverage/findings.md β "Can-fail proofs".AssertTrue(cond, msg) (returns the bool, records a failure), SetResultSuccess(), SetResultFailure("why"). The failure string appears verbatim in junit.xml and autotest.log.Print / PrintFormat, not the global ones β the shadowed versions route through the autotest printer into autotest.log. Global Print only reaches console.log.#ifdef WORKBENCH a test class. Guarded tests do not exist in the retail client, which is what CI runs. Test code is inert without -autotest at runtime, which is the real safety guarantee.Worlds/MP/OVT_Campaign_Test.ent loads OVT_OverthrowGameMode plus managers via its default.layer, so OVT_Global accessors work β but the campaign is not started. Do not start it by hand: override RequiresStartedCampaign() on the suite and let OVT_TEST_SuiteBase do it. That sequence is non-idempotent, must select the difficulty preset by name ("Test World" β index 0 is Easy), and has to close the start menu; one guarded implementation exists so every campaign-tier suite starts identically.OVT_Global is fine β its statics were measured against the live game mode across the harness's three world loads and agreed 100% of the time (the engine nulls the weak statics on destruction). OVT_TEST_SuiteBase.ResolveManager(typename) finds a component on the live game mode and exists as defensive documentation if that ever changes. Tier A may use neither β it has no game mode at all.>= 1, never a magic count.new does not apply [Attribute()] defvalues β a hand-built config object starts fully zeroed, which silently breaks any field whose declared default is not zero (an "unset" sentinel of -1, a multiplier of 1). Set every field a Tier A case depends on, explicitly.tools/compile-check.sh first, then tools/run-tests.sh <YourSuite>, then the Fast or All group.Detailed documentation organized by concern:
tools/compile-check.sh is the build-verification equivalent (compile only, no artifacts)tools/run-tests.sh runs Reforger's shipped autotest framework against Overthrow: 30 assertions over logic, manager init, started-campaign state and same-session persistence. JIP/multiplayer, UI, performance and save/reload are uncovered and stay manualPattern: Start here for quick reference, dive into resource files for detailed procedures.