Skip to content

Vulkan backend with TAA and GTAO - #110

Open
NightHammer1000 wants to merge 297 commits into
StratumServer:mainfrom
KillerPixelCrew:feat/vulkan-taa
Open

NightHammer1000 wants to merge 297 commits into
StratumServer:mainfrom
KillerPixelCrew:feat/vulkan-taa

Conversation

@NightHammer1000

@NightHammer1000 NightHammer1000 commented Sep 17, 2026 •

Copy link
Copy Markdown

Follow-up to #69. This draft delivers the native Vulkan renderer, native shaders, TAA, and GTAO. Vulkan is opt-in through Renderer in optimum.json; OpenGL remains the default. New upscaler integrations, frame generation, Vulkan-to-DX12 interop, and OpenGL interop for those features are outside this PR. The native final blit supports Optimum's existing FSR 1 render-scale option.

Delivered

  • Vulkan 1.3 rendering and presentation with explicit synchronization, asynchronous uploads, a streaming frame graph, shader and pipeline caches, and native draws for terrain, entities, particles, sky, GUI, and the post chain. The client-facing stated-state adapter remains for render API and mod compatibility.
  • 50 native shader programs maintained as one stage-selected source per program, with offline SPIR-V packaging and a runtime rewriter for mod shaders.
  • TAA with motion for world and hand geometry, history rejection, anti-flicker weighting, and sharpening after final composition. Bloom and god rays consume the unsharpened scene.
  • GTAO composed before TAA, selectable alongside vanilla SSAO. The per-block optimumAoThin class handles thin geometry without treating snow layers as thin.
  • Separate scene and UI images, premultiplied UI composition, frame/submit/present identity, CPU phase timing, and per-pass GPU timing. These are renderer foundations; vendor latency and frame-generation backends are outside this PR.
  • ShadowCasterSplit was dropped after a measured far-cascade regression. Shadow draw batching and culling remain follow-up work.

Refactor and scope

The current diff against StratumServer:main has 371 changed paths, down from the earlier 719-file state. The maintained Vulkan tree has 76 shader source files (73 GLSL and 3 compute), 83 renderer C# files, and 41 test C# files, of which 34 are test suites. Redundant raw integration harnesses were replaced with focused device regressions, including direct RGBA16F checks for terrain, object, entity, liquid, particle, and sky motion. Research notes were removed; docs/vulkan.md is the single tracked Vulkan guide. The optional scripted headless capture harness is on the separate codex/headless-capture branch; no PR has been opened for it.

Verification

  • Current Release solution: 1,554 passed, 34 existing skips, 0 failed. This includes all 504 Vulkan tests on the explicitly selected Intel UHD Graphics 770 (Xe-LP), 744 core tests, and 306 launcher, CLI, bootstrap, and installer tests.
  • All 25 consolidated motion contracts passed on the NVIDIA GeForce RTX 4070 Laptop GPU. The preceding full 556-test renderer suite passed on NVIDIA before the test-only motion replacements. Both GPU runs loaded Khronos validation with synchronization and best-practices checks.
  • Strict donor validation: 93 source patches and 68 Cecil patches applied, and all 43 runtime patches compiled against the pinned game donor.
  • All 286 packaged SPIR-V module paths and bytes match the saved pre-refactor baseline. The two earlier Intel regressions, transient post-chain aliasing and TAA history/output reuse, are covered by the current passing GPU suite.
  • The branch is zero commits behind StratumServer:main; there are no unresolved merge conflicts.

@NightHammer1000

NightHammer1000 commented Sep 17, 2026 •

Copy link
Copy Markdown
Author

This approach fundamentally differs from the previous Vulkan implementation that was emulating OpenGL Calls.

Its the proper way to do this to actually get Vulkans full potential here.
Only Issue is that it cant support Mods that make GL Calls on their own.
So a Mod API surface is a future point to think about.

The UI separation and frame marking were backported from my DLSS and DLSS-FG branch as foundational systems here, as I will need them later on for upscaling, frame generation and VR

…reports 43 runtime patches applied and exact donors compiled

The installed-launcher path decompiles the user's own VSEssentials.dll and
VSSurvivalMod.dll, applies patches/runtime/**, compiles that, and lets Cecil
transplant the bodies named in Optimum.Patcher/mod-patcher.cs. Every TAA P3/P4
mover had a mod-patcher Methods entry but no runtime donor, so installed players
would have got vanilla bodies: no motion vectors, camera-fallback ghosting, with
all fork tests green.

21 new donor patches, generated against a pristine .build/runtime-donors
decompile and matching the fork sources member for member:
- VSEssentials: EntityPlayerShapeRenderer (hand window), ModSystemFpHands
  (AnimationPrev UBO), ModSystemRenderFallingBlocksFast (per-entity previous
  matrix, one window around the loop).
- VSSurvivalMod: Quern, HelveHammer, Fruitpress, Resonator, Bloomery, Forge,
  Firepit, PotInFirepit (pot + lid), EchoChamber.
- Mechanics: MechBlockRenderer (NoteDevice, WriteInstance), MechNetworkRenderer
  (ApplyPassUniforms + window), AngledCageGear, AngledGears, GenericMechBlock,
  Transmission, Clutch, CreativeRotor, Pulverizer (OptimumInstanceMotion layout,
  instance counts).
EntityItemRenderer and EntityShapeRenderer donors regenerated to carry their
RenderItem/DoRender3DOpaque motion additions on top of what they already had.

Tests: new Optimum.Tests/taa-runtime-donor-coverage-tests.cs pins every fork
patch that adds a motion call to a runtime donor that adds the same calls, and
fails when a new instrumented fork patch has no donor mapping.
mod-patcher-manifest-consistency-tests.cs gains a Methods cross-check: every
transplanted method's declaring type must have a runtime patch or a sources/**
overlay, with the two FluffyClouds (Vulkan, not TAA) entries listed as known
gaps so a new one fails.

Verified: scripts/extract-patches.sh clean; scripts/check-patches.sh reports
"Runtime patches: 43 applied and exact donors compiled" (was 22); both donor
projects compile; dotnet build VintageStory.slnx -c Release; dotnet test
Optimum.Tests -c Release 951 passed; dotnet test Optimum.Render.Vulkan.Tests
330 passed. Donor coverage guards checked negatively by removing one patch.
No fork, lib, shader or manifest behaviour changed.
…all tests green

Adds the tooling P5 needs before the acceptance matrix can be run; no renderer
behaviour changes and nothing new runs on the GPU.

- scripts/dev/perf-capture.sh: launches through scripts/dev/run-client.sh with
  RENDERER=<backend>, optionally rewrites "Taa" in optimum.json, waits for the
  "[Client Chat] Welcome" line plus an 8 s warm-up, records a 30 s window, closes
  through scripts/dev/kill-client.sh and prints mean / 1% low / worst frame time
  plus a summary.csv. It re-reads the renderer out of the client log and refuses
  to report numbers on a silent OpenGL fallback (rule 1). It never pattern-kills.
- build/VintagestoryLib ClientMain: OptimumLogFrameTime writes one line per second
  ("[Optimum] fps window= frames= mean= min= max= p99=") when OPTIMUM_FPS_LOG names
  a file. Off otherwise - one null check per frame - so TAA off is unchanged.
  Backend-neutral, unlike OPTIMUM_VULKAN_STATS, which only exists on Vulkan.
  Members and the method are listed in Optimum.Patcher/Program.cs; the caller
  MainRenderLoop is an existing transplant target.
- scripts/dev/luma-diff.py: the still-frame luminance diff from the parity skill
  (centre 60% crop, --median over seven pairs), so the doc can name a real command.
- docs/taa-acceptance.md: the P5 matrix as a checklist - 17 visual rows plus
  byte-identical-off, performance and memory, each with scene, exact commands,
  pass criterion and the measurement to record; preconditions (creative, wind
  stilled, storms off, noon, clear sky); the P4 items carried into P5. docs/ is
  git-ignored, so .gitignore gains a docs/* + negation for this one deliverable.
- Optimum.Tests/taa-acceptance-harness-coverage-tests.cs: the doc names every
  matrix row TAA-PLAN.md lists, each row carries its four parts, the script drives
  the dev scripts and confirms the renderer, and the patcher ships the log members.

Verified: bash scripts/extract-patches.sh; bash scripts/check-patches.sh
(93 applied, 64 cecil, 0 pending/conflict; runtime patches 22 applied and exact
donors compiled); dotnet build VintageStory.slnx -c Release (0 errors);
dotnet test Optimum.Tests -c Release (929 passed, 34 skipped, 0 failed);
dotnet test Optimum.Render.Vulkan.Tests (330 passed, 0 failed).
NOT verified: the game was never launched (task prohibition), so the capture
script has not been run end to end against a live client.
Merges the four P5 worktrees into feat/taa: sharpen pass + mip bias,
settings rows / scanner rules / shader packaging, runtime donors for the
P3/P4 mod movers, and the acceptance + performance harness.

Conflict resolution:
- Optimum.Launcher/ShaderCompatibilityScanner.cs: both branches extended the
  externalMotionShader veto. Kept the superset - taa-resolve, taa-debug,
  taa-skymotion and taa-sharpen named individually, plus the taa- prefix rule,
  the FSR pair the sharpen shares its lobe maths with, vertexwarp.vsh and the
  whole shaderinclude directory. The duplicate taa-resolve/taa-debug lines the
  two sides each added were folded into one.
- Optimum.Patcher/Program.cs merged both sides (sharpen members and program,
  the three settings callbacks, the fps-log fields and OptimumLogFrameTime).
- build/ and the API fork are per-worktree and git-ignored, so the winning
  sources were copied into the main tree before re-extracting: ChunkRenderer,
  ClientPlatformWindows, ShaderPrograms, ShaderRegistry and
  VintagestoryApi/Config/OptimumConfig.cs from the sharpen branch,
  GuiCompositeSettings from the settings branch, ClientMain from the harness
  branch. The mod-fork donor branch touches only patches/runtime/**, which
  extract does not regenerate.

Merge gap fixed: the Windows packager's explicit "what a release contains"
list was written before the sharpen shader existed, so taa-sharpen.vsh/.fsh
shipped only through the wildcard overlay and nothing asserted them.
Added to scripts/package.ps1 and to both lists in
taa-settings-coverage-tests.cs.

Verified: scripts/extract-patches.sh clean (157 patches, 0 stale);
scripts/check-patches.sh 0 pending / 0 conflict, runtime patches 43 applied
and exact donors compiled; dotnet build VintageStory.slnx -c Release 0 errors;
Optimum.Tests 987 passed / 34 skipped; Optimum.Render.Vulkan.Tests 333 passed;
Optimum.Launcher.Tests 32 passed. Game not launched.
…, NaN LOD-bias cache

Adversarial review of the P5 range (a04f409..f1a6300). Two defects found and fixed;
everything else in the range was checked and held (see below).

1. The TAA mip-bias row did not apply live to the passes it exists for.
   chunkopaque and chunktopsoil sample the block atlas through sampler OBJECTS,
   and a bound sampler object overrides the texture object's parameters on that
   unit - LOD bias included. The bias was written into those samplers only in
   ShaderRegistry.loadRegisteredShaderPrograms, i.e. once per shader load, while
   ChunkRenderer.ApplyOptimumTextureLodBias re-applied the atlas TexParameter
   every frame. So dragging the slider moved the mip selection of liquid,
   transparent and shadow terrain and left opaque terrain and topsoil on the bias
   they were compiled with - an inconsistent mip choice across passes of the same
   atlas, with the GUI tooltip and onOptimumTaaMipBiasChanged both promising
   "applies immediately".
   Fix: ShaderRegistry.ApplyOptimumTerrainSamplerLodBias(float) is now the single
   place those four samplers are written, called from the shader load as before
   and from ChunkRenderer.SetOptimumTextureLodBias, which already only fires when
   OptimumConfig.EffectiveTerrainLodBias actually moves. Both new members are in
   Optimum.Patcher/Program.cs.

2. ChunkRenderer.optimumTextureLodBias started at 0f, not NaN, so the first
   OnBeforeRenderOpaque of a TAA-off native-scale session read "0 wanted, cache
   not NaN" and wrote an explicit LOD bias of 0 over the driver default on every
   atlas - the one call the code comments and the coverage test say that
   configuration never makes. With fix 1 that would have reached every terrain
   sampler too. Initialised to float.NaN.

Regression tests:
- Optimum.Render.Vulkan.Tests/VulkanDeviceIntegrationTests
  ALodBiasWrittenToAnAlreadyBoundSamplerChangesTheMipTheGpuReads: a 4x4 mip chain
  with one flat colour per level, drawn at one texel per pixel so lambda is 0;
  SetSamplerParameter on the already-bound sampler moves the readback from level
  0 to level 1 and back. This is the backend half of the claim - the descriptor
  set has to key on the resolved VkSampler, not on the sampler id.
- Optimum.Tests/taa-sharpen-coverage-tests: ALiveMipBiasChangeReachesThe
  TerrainSamplerObjectsAsWell and TheCachedLodBiasStartsAtNanSoTaaOffTouchesNothing.
- fsr-pipeline-coverage-tests updated for the extracted sampler helper.

Checked and NOT changed: TAA-off byte-identity (the FXAA/Luma else branch is
vanilla; index 21 and the sharpen pass are both gated); the no-double-sharpening
rule (one shared OptimumFsrBlitActive(), term-for-term the old inline condition);
sharpness 0 (early-out before the first ring tap, GPU-proven bit-identical);
scanner verdicts (taa- prefix plus the whole shaderincludes directory, external
sources only); packaging (all five packagers, the Makefile wildcard-plus-verify,
package.ps1's required-file list); runtime donors (43 applied, exact donors
compiled, marker parity per fork patch); patcher entries; and the translation gate
- taa-sharpen is a program pair in sources/shaders, so ShaderCorpus.ProgramNames
picks it up and EveryVanillaProgramTranslatesToSpirv covers every variant row.

Verified: extract-patches + check-patches (93 applied, 64 cecil, 0 pending; 43
runtime patches, exact donors compiled), Release build, 989 source tests, 334 GPU
tests, 32 launcher tests. Not verified in game - no deploy, no client run.
What P5 landed (sharpen pass, no-double-sharpening rule, mip bias, settings rows,
scanner rules, packaging verification, the installed-runtime donors that close P3
finding (f) / P4 finding (t), and the acceptance + performance harness), the exact
vs fallback table updated for it, findings (ab)-(af) from the adversarial review,
and the full "still owed in the game" list.

The acceptance matrix has not been run - no phase of P5 deployed or launched the
client - so P5 is not done by rule 3 and the default-on decision is left to the
user with docs/taa-acceptance.md as the gate. TAA stays opt-in.
perf-capture and rule 1 both key on the '[Optimum] <backend> renderer' line; an explicit OpenGL choice, and a Vulkan start without an advisory, printed nothing.
docs/temporal-frame-contract.md is now the frozen, versioned (v1) specification
every temporal consumer is written against: the in-house TAA resolve today,
FSR 3.1 / XeSS 2 / DLSS next, frame generation and ray reconstruction after that.

It specifies the per-frame input record member by member (type, units, coordinate
convention, the point in the frame after which each value is this frame's), every
resource with format, resolution, sampler state and channel semantics - the motion
attachment's rg/b/a including the writer-depth validity tolerance and its
half-float rationale, the history colour/glow/linear-depth slots and why their
filters differ, the sharpen target - the jitter definition and Halton sequence,
the reset reasons with their triggers, the per-class exact/fallback/reactive
status consolidated from P3-P5, and the adapter formulas for the three vendor
upscalers (motion vector scale and sign, jitter sign, depth convention,
reactive/transparency mask mapping, exposure, camera constants). Native handles,
extension negotiation, presentation lifetime and ray-reconstruction guides are
explicitly reserved for the vendor plan.

Optimum.Tests/temporal-contract-tests.cs is the tripwire: it pins the public
surface of IOptimumTemporalContext and OptimumTemporalFrame by reflection against
a checked-in list, and pins the conventions the document states as fact - the
shear formula in both implementations and numerically, the motion-vector scale
and sign per adapter, the writer-depth tolerance expression in taa-resolve.fsh,
the resolve's full input set, the history/sharpen slot indices and parity rule,
and the attachment formats and sampler state on both backends. Every failure
message names the document, and the surface assertion prints the actual list so
the new one is the failure message.

Verified: extract-patches + check-patches (93 applied, 64 cecil, 0 pending; 43
runtime donors exact), dotnet build VintageStory.slnx -c Release, dotnet test
Optimum.Tests -c Release (1007 passed, 34 skipped). Negative control run: removing
one member from the checked-in surface list fails TheTemporalContextSurfaceIsFrozen
with the document named. No game run - this phase changes no rendering code.
…itch reset, ChunkRenderer motion windows exception-safe

- ClientPlatformWindows: both frame-buffer setup paths (device and GL) now read
  `!optimumTaaDisabled && OptimumConfig.EffectiveTaa`, so a platform that already
  failed the TAA allocation never retries it on a later rebuild. EffectiveTaa
  alone already goes false (DisableOptimumTaa always calls DisableTaaAtRuntime),
  so this is a belt-and-braces guard on the per-instance flag; the stale doc
  comment on DisableOptimumTaa that claimed EffectiveTaa stays true was corrected.
- GuiCompositeSettings.onOptimumTaaChanged: the explicit-disable bail-out now
  resets the already-flipped switch to EffectiveTaa before returning, so the UI
  cannot show TAA on while it is off for the session.
- ChunkRenderer.RenderOpaque and RenderAfterOIT: the motion-vector windows are
  wrapped in try/finally (matching RenderLiquidMotion), so a throwing shader
  setup or pool draw can no longer leak the expanded draw-buffer mask into every
  later draw and block every later window. Draw order and state calls unchanged;
  GlPopMatrix left where it was.

Verified: dotnet build VintageStory.slnx -c Release (0 warnings, 0 errors);
scripts/extract-patches.sh + scripts/check-patches.sh (0 pending, 0 conflict,
157 patches, 43 runtime donors exact); dotnet test Optimum.Tests -c Release
(1010 passed, 0 failed, 34 skipped), including the three new coverage tests.
All changed methods were already listed in Optimum.Patcher/Program.cs.
…pply inside the window; runtime donors mirrored

Verified: dotnet build VintageStory.slnx -c Release (0 errors);
extract-patches + check-patches (93 applied, 0 pending, 0 conflict;
43 runtime patches applied and exact donors compiled);
dotnet test Optimum.Tests -c Release (1011 passed, 0 failed, 34 skipped).
Test-only changes, no production code touched.

- fsr-pipeline / taa-sharpen: "if (textureLodBias == 0f)" now has to be the
  restore path it actually is - the float.NaN cache, the !IsNaN guard,
  SetOptimumTextureLodBias(0f) and the cache reset, all asserted inside that
  branch (brace-matched, so the nonzero path below cannot satisfy them). The
  sampler entry point is asserted to apply the raw value: no "!= 0f" inside
  ApplyOptimumTerrainSamplerLodBias, all four samplers written with `bias`.
  Read ShaderRegistry.cs:315-343 first: the sampler path does NOT skip zero, so
  there is no production gap here.
- New PatchMethodScopes: locates a patch's added lines inside the tree it was
  applied to and attributes them to the enclosing method (brace scanner that
  skips comments and string literals). Covered by patch-method-scopes-tests.cs.
- mod-patcher-manifest-consistency: EveryTransplantedMethodHasARuntimeDonor is
  per METHOD now, not per declaring type. A type whose donor patch touches only
  some other method is reported separately (KnownUnchangedTransplants, seeded
  with ChunkMapLayer::loadFromChunkPixels, whose body really is unchanged in
  both trees). KnownDonorGaps keeps its meaning. Degrades to the old per-type
  check when neither .build/runtime-donors nor the fork is on disk.
- taa-runtime-donor-coverage: marker sets are compared per method as well as
  file-wide, so a donor that writes the same markers into a different body no
  longer passes.
- taa-terrain-motion: the prepass test reads through ReadPatchedOrSource like
  the rest of the file (it would have thrown in a clean clone); the vanilla
  projection line, which sits outside every hunk, is asserted against the
  decompiled tree only when that is checked out.

Verified: dotnet test Optimum.Tests -c Release - 1015 passed, 0 failed, 34
skipped. Negative check on real data: moving the Optimum block of
QuernTopRenderer out of OnRenderFrame in .build/runtime-donors made both
strengthened tests fail with the right messages (donor coverage listed the
three misplaced markers; the manifest test named
QuernTopRenderer::OnRenderFrame); donor restored afterwards.
…t masking, dead UBO buffer

- CreateInstance: enable VK_EXT_validation_features whenever the
  ValidationFeaturesEXT struct is chained into pNext; the feature list is now
  parsed (ParseValidationFeatures) before the extension array is marshalled so
  both decisions use the same condition. A chained struct with the extension off
  was silently ignored.
- ProgramInterfaceLayout: fragment output arrays with only a constant element
  stored to now mark just that element's location written
  (TryGetWrittenFragmentOutputElements); dynamic indices, whole-array and
  swizzled stores keep the full span, and an index on a non-array output still
  means the whole variable. PipelineCache no longer leaves colour writes on for
  an attachment the shader never writes (undefined data on Vulkan).
- VulkanDevice: ResolveValidationLogPath accepts a Windows path too, and the
  default-path field now precedes the field that reads it (static initialiser
  order made the bare OPTIMUM_VULKAN_VALIDATION=1 log path null).
- MirrorValidationMessage also swallows UnauthorizedAccessException,
  NotSupportedException and ArgumentException; a diagnostic write must not abort
  Initialize.
- ClientUniformBuffer: removed the per-UBO VulkanBuffer, SyncBuffer, the
  descriptor Release and the deferred disposal. Verified dead: no path binds
  ubo.Buffer - the ring-exhausted path allocates its own transient copy - and
  SyncBuffer had no caller anywhere in the tree.
- TaaResolveTests: all eight LoadProgram sites are 'using' now, so the programs
  are destroyed before the context.

Verified: dotnet build Optimum.Render.Vulkan.Tests -c Release (0 warnings,
0 errors) and dotnet test Optimum.Render.Vulkan.Tests -c Release on the GPU with
validation layers on: 349 passed, 0 failed, 0 skipped (ShaderTranslationTests
and the pipeline/attachment tests included). Not run here: the game, make deploy,
extract/check-patches (ignored trees absent in this worktree).
…f pairing, --vsync validation, doc corrections

Applies the review findings scoped to the launcher scanner, the Makefile deploy,
the dev scripts and the TAA docs.

- ShaderCompatibilityScanner.NormalizeShaderPath: the stage-extension filter
  (.fsh/.vsh/.gsh) now applies to shaders/ only. ShaderRegistry loads every
  shaderinclude regardless of extension and vanilla ships five .ash includes, so
  an external .ash override was invisible and TAA stayed on with a replaced
  helper. Directory entries (no extension) are still rejected.
- Makefile deploy: shader and shaderinclude copies are per-file with
  `cp -f ... || exit 1` instead of `find -exec ... \;` / a bare wildcard cp
  (find returns 0 when an individual cp fails), and the completeness check
  compares CONTENT with `cmp -s` instead of `[ -f ]`, which passed against an
  unchanged vanilla file of the same name. Both the VANILLA_DIR and the
  INSTALL_DIR block.
- scripts/dev/luma-diff.py: --median now pairs non-overlapping files
  (files[::2] with files[1::2]) instead of every adjacent pair, errors on an odd
  count, and the docstring plus the docs/taa-acceptance.md invocation say to pass
  files in pair order (a flat `shots/*.png` glob sorted a1..a7 before b1..b7).
- scripts/dev/perf-capture.sh: --vsync is validated like --taa ("", on, off),
  exit 2 to stderr; "--vsync 0" or a typo silently meant vsync ON.
- docs/taa-acceptance.md P2: the slot-21 sharpen target (OptimumTaaSharpenIndex
  = 21, RGBA8 the size of Primary, ~15.8 MiB at 1080p) added to the memory
  budget; total ~86.9 MiB.
- docs/temporal-frame-contract.md: escaped |CameraPosDelta| in the Teleport row
  (it read as extra table cells); new section 6.1 declaring cloud pixels
  UNSUPPORTED for external motion consumers (mv is camera-rotation-only, reject
  via the reactive mask; a real vector needs prev cloudOffset + the ray-marched
  hit from cloudvolumetric.fsh, future work). No shader changed.
- TAA-PLAN.md: finding (n) carries the same contract term; finding (t) rewritten
  - the eight mover donors landed in P5 (c897e23, merged 0120422) under
  patches/runtime/VSSurvivalMod/Vintagestory/GameContent/ (MechNetworkRenderer
  under .../Mechanics/), guarded by Optimum.Tests/taa-runtime-donor-coverage-tests.cs
  and mod-patcher-manifest-consistency-tests.cs; the two P4 status paragraphs no
  longer claim "not verified in game on either backend" - the entity
  motion-writer gate was verified on Vulkan and OpenGL (2830577) and P5's
  in-game run (7b0168d, deployed c9758ce+5b952da) drove both backends, while the
  per-class mover behaviour and the 18-row acceptance matrix stay unverified.

Verified: dotnet test Optimum.Launcher.Tests -c Release -> 34 passed, 1 failed
(SolutionIntegrityTests.EverySolutionProjectPathExists, pre-existing and only
because the git-ignored build/ tree is absent in this worktree); the scanner
filter alone is 18/18 green. `make -n deploy` exits 0 and the rewritten copy and
cmp snippets were run against a temp directory: a clean copy passes, a tampered
destination and a deleted destination both fail with the new message.
`bash -n scripts/dev/perf-capture.sh` clean and
`perf-capture.sh --renderer vulkan --vsync bogus` exits 2 before any config
write. luma-diff.py run on generated 64x64 PNGs: 2 pairs -> median 5.000, odd
count -> SystemExit. Not run here: extract-patches/check-patches and the game
(no ignored trees in this worktree).
…ged Makefile

The scripts/docs stage replaced the deploy shader copy (find -exec cp, then a
bare [ -f "$d" ] completeness check) with a per-file cp loop that aborts on
failure and a cmp -s content comparison. MakeDeployCopiesEveryShaderAndFails-
WhenOneDoesNotArrive still pinned the old literal, so the merge of the two
branches broke it - a semantic conflict, no functional regression.

The assertions now pin the stronger contract: both deploy paths loop over the
whole sources/shaders and sources/shaderincludes directories (not *.fsh plus
*.vsh), each cp is '|| exit 1', and both completeness checks compare content
with cmp -s rather than existence, with a DoesNotContain on the old [ -f ] form.

Verified: dotnet build VintageStory.slnx -c Release 0 errors; extract-patches
wrote 157 patches with no tree change; check-patches 0 pending / 0 conflict,
43 runtime donors exact; Optimum.Tests 1015 passed 0 failed 34 skipped;
Optimum.Launcher.Tests 35 passed 0 failed; Optimum.Render.Vulkan.Tests 349
passed 0 failed on a real GPU with validation layers on.
…orrect the sharpen budget

Adversarial review of c60a4cc..b869ddd. Two defects found and fixed; the rest
of the diff verified as described.

1. VulkanContext.CreateInstance: finding 1 added VK_EXT_validation_features to
   the instance extension list without checking that the layer advertises it.
   The extension is deprecated in favour of VK_EXT_layer_settings; naming one
   the layer does not have fails vkCreateInstance with ErrorExtensionNotPresent,
   TryCreate then reports no usable device and the bootstrap falls back to
   OpenGL without a word - so OPTIMUM_VULKAN_VALIDATION_FEATURES could have
   turned into "Vulkan silently stopped working" (rule 1). Chaining is now
   gated on LayerAdvertisesExtension(); when the extension is absent the
   behaviour is what it was before the fix (struct not chained, instance comes
   up) instead of no Vulkan at all.
   Test: ValidationFeaturesTests.TheFeaturesExtensionIsCheckedAgainstTheLayer -
   an invented extension name and an uninstalled layer both answer false.

2. docs/taa-acceptance.md P2: the slot-21 sharpen target was described as
   RGBA8 while both allocation paths use EnumTextureInternalFormat.Rgba16f
   (ClientPlatformWindows.cs:1834 device, :2426 GL). The 15.8 MiB figure and
   the ~86.9 MiB total were right for RGBA16F; only the format word was wrong.

Also corrected the now-stale "Same order as DeleteUniformBuffer" comment on the
ring-exhausted UBO path, which finding 5 left pointing at code it deleted.

Verified, no change needed: hunk arithmetic of both hand-written runtime donor
patches (counted: 115,10 -> 115,27 and 73,7 -> 73,29); bash scripts/check-patches.sh
reports 93 applied / 0 pending / 0 conflict and 43 runtime patches with exact
donors compiled; ChunkRenderer RenderOpaque and RenderAfterOIT keep their draw
order, profiler marks and GL state with the window closed in a finally;
GuiCompositeSettings' composer field is the optimum composer (assigned at 1767)
and "optTaa" matches the AddSwitch row; ProgramInterfaceLayout's store regex only
widens the matched language (superset of the old one, so no output can newly be
masked off) and no shader in the deployed corpus declares an array fragment
output at all; no reference to ClientUniformBuffer.Buffer/SyncBuffer survives;
Makefile loops abort on a failed cp and the completeness check now compares
content.

Tests: dotnet test Optimum.Tests -c Release -> 1015 passed, 34 skipped, 0 failed.
dotnet test Optimum.Launcher.Tests -c Release -> 35 passed, 0 failed.
dotnet test Optimum.Render.Vulkan.Tests -c Release -> 350 passed, 0 skipped,
0 failed (349 before this commit's new test).
…econdary worktree offline

Mirrors scripts/bootstrap.sh steps 7 and 8 on top of the main checkout's .baseline and links
.vanilla, _ref, .baseline and .build. Verified in a throwaway worktree: 157 patches applied,
build/VintagestoryLib byte-identical to the main checkout, dotnet build VintageStory.slnx -c
Release 0 errors, Optimum.Tests 1015 passed once the lib output exists.
…kout after a merge, keeps bin/obj

Refuses when patches/ or sources/ is dirty or when extract-patches.sh would write anything
(unextracted edits in build/ or a fork). Verified in the main checkout: build/VintagestoryLib,
VintagestoryApi and VSEssentials hash identical before and after, VintagestoryLib.dll kept.
…ng after a merge

Verified: plain --in-place passes on the clean main checkout, and --discard-build-edits
re-materialises 157 patches with no working-tree changes.
…rt dispatch verifier; donor ClientPlatformWindows unsealed with 4 public virtuals

Verified: dotnet build VintageStory.slnx -c Release 0 errors; Optimum.Tests 13 new tests pass,
991/991 pass excluding the runtime-donor classes (their shared .build tree was clobbered by
check-patches.sh through the worktree symlink, not by this change); patcher run as in make patch-il
into bin/: 1 type unsealed, 4 methods virtualized, 257/257 transplanted, dispatch verifier ok
(24 callvirt sites, 0 call/ldftn); check-vanilla-compat.sh ok with no allowlist change;
extract-patches 0 conflict 0 pending.
…ounters, OPTIMUM_FPS_LOG stddev, pacing-gate.sh

Verified: dotnet build VintageStory.slnx -c Release 0 errors 0 warnings;
dotnet test Optimum.Render.Vulkan.Tests 360 passed 1 skipped 0 failed (incl.
TextureUploadInsideAFrameBlocksOnTheUploadSiteUntilPhase1B on the GPU);
Optimum.Tests: new and touched classes 43/43 green, full suite 998 passed and
25 failed, all in runtime-donor tests (.build/runtime-donors clobbered through
the worktree symlink by check-patches), none in files this commit touches;
extract-patches 0 conflicts 0 pending; pacing-gate.sh --self-test 12/12.
Not run in game.
…nned sync-hazard ledger; OPTIMUM_VULKAN_POISON

Every GPU test gets its context or device from GpuTest (EnableValidation, features sync,best, overridable via OPTIMUM_TEST_VALIDATION_FEATURES). ValidationAssert.NoSyncHazards runs wherever NoErrors does and fails on SYNC-* messages not pinned in KnownSyncHazards; SyncHazardLedgerTests runs last and fails when a pinned entry no longer occurs. Debug messages now carry the layer message id. Poison mode fills fresh images (NaN float, magenta UNORM/SRGB, 0xDEADBEEF int, 0.5 depth) and host-visible buffers (0xDEADBEEF).

Verified: dotnet test Optimum.Render.Vulkan.Tests 373 passed, 0 failed, 1 skipped (no NV checkpoints); zero SYNC messages over 105 asserting tests, pinned list empty; positive control reports SYNC-HAZARD-WRITE-AFTER-WRITE/READ-AFTER-WRITE/WRITE-AFTER-READ. Solution build 0 errors. Optimum.Tests 994 passed, 25 failed, all runtime-donor manifest tests whose inputs (patches, Optimum.Patcher, forks, sources) are unchanged from 15da683.
…rity-capture.sh, acceptance docs

OPTIMUM_PARITY_DUMP=<abs dir> OPTIMUM_PARITY_FRAME=<n>: ClientPlatformWindows counts frames
rendered while the player is in the world (BlocksReceivedAndLoaded) and on frame n, after the
post chain and final blit and before Present/SwapBuffers, dumps every attachment of every
framebuffer slot once (shared textures once), then logs "[Optimum] parity dump: <count>
attachments -> <dir>". Env unset: one static bool check per frame, no GL/Vulkan call.
Shared writer OptimumParityDump (API): one file-name format, PPM+alpha PGM for 8-bit unorm,
little-endian PFM (negative scale) for float and depth, rows bottom-up in GL order.
GL side: glGetTexImage (pack buffer unbound and restored). Vulkan side: one seam member
ReadTextureForParity, readback shared with TextureDump (depth aspect now handled), requested
GL token recorded on VulkanTexture so promoted formats pair by name.
New: scripts/dev/ssim.py (+ --self-test), scripts/dev/parity-capture.sh (not run),
docs/parity-allowlist.md, docs/vulkan-acceptance.md (.gitignore exceptions).

Verified: dotnet build VintageStory.slnx -c Release 0 errors 0 warnings;
Optimum.Render.Vulkan.Tests 350 passed, 1 skipped (GpuCheckpointTests), 0 failed, incl.
ParityDumpTests (RGBA8/RGBA16F/R32F/depth patterns decode back in GL row order);
Optimum.Tests 998 passed, 25 failed, all in ModPatcherManifestConsistencyTests and
TaaRuntimeDonorCoverageTests (patches/runtime vs fork patches, untouched here);
ParityDumpCoverageTests 8/8 passed; ssim.py --self-test ok. extract-patches ran; check-patches
exits 1 listing only VSSurvivalMod/VSEssentials "requires a patch refresh" lines.
Not verified in game: no launch or deploy in this stage.
…cher and parity

Resolved source: parity dump members kept, followed by the unsealed class and the four
public virtuals from the patcher stage. extract-patches: 157 patches, 0 stale;
check-patches: 93 applied, 64 cecil, 0 pending, 0 conflict; runtime patches 43 applied.
Integration fixes only:
- VulkanDevice.DumpRequestedTextures/ReadBackLevel0: parity's shared readback kept,
  with the stats stage's counted waits (VulkanStats.WaitDeviceIdle, WaitSite.Readback).
- PacingStatsTests and ParityDumpTests get their device from GpuTest (sync,best
  validation, AssertClean) instead of a private new VulkanDevice.
- parity-dump-coverage-tests searches the now-virtual SetupDefaultFrameBuffers signature.

Verified: dotnet build VintageStory.slnx -c Release 0 errors; Optimum.Tests 1048 passed,
0 failed, 34 skipped; Optimum.Render.Vulkan.Tests 385 passed, 0 failed, 1 skipped
(GpuCheckpointTests); check-patches 0 pending 0 conflict, runtime patches 43 applied;
pacing-gate.sh --self-test 12/12; ssim.py --self-test ok. Not run in game.
… and the helper check

Adversarial review of Phase 0 (15da683..4d1089f). Fixes:
- VulkanStats: new wait site queue_submit. FrameSlot.EndFrameAndSubmit waited on
  QueueLock, which a worker's SubmitAndWait holds through its fence wait, and that
  render-thread stall was not counted anywhere.
- parity-capture.sh: set -euo pipefail with errexit-safe greps, and cleanup polls for
  the client to exit before restoring optimum.json (kill-client.sh does not wait).
  The sleep after the client exits in wait_for is gone.
- pacing-gate.sh and perf-capture.sh: set -euo pipefail. perf-capture's renderer
  grep can no longer end the script with the client still running.
- vulkan-test-validation coverage: the helper-only check now also catches
  target-typed new() and EnableValidation = false.

Verified: extract-patches (157 patches, no drift) and check-patches (0 pending, 0 conflict);
dotnet build Release 0 errors; Optimum.Tests 1056 passed / 34 skipped / 0 failed;
Optimum.Render.Vulkan.Tests 386 passed / 1 skipped / 0 failed, including the new
AFrameSubmitIsCountedAtTheQueueSubmitSite; pacing-gate --self-test 12 cases; ssim.py
--self-test ok; the Cecil lib patch run into a scratch output passed: 257/257 methods,
1 type unsealed, 4 methods virtualized, verifier ok with 24 callvirt and 0 call.
Nothing was deployed or run in game.
check-patches.sh (prepare-runtime-donors.sh) wipes and re-decompiles .build/runtime-donors; through
the symlink the Phase 0 stages destroyed the main checkout's donor tree and 25 runtime-donor tests
failed everywhere until check-patches ran in the main checkout. Verified in a throwaway worktree:
.build is a directory (own inode, reflink copy), .vanilla/_ref/.baseline stay links, 157 patches
applied; prepare-runtime-donors.sh writes only under .build.
… world loads

A client that crashed at window creation (GLXBadFBConfig after an NVIDIA driver update without a
reboot) left perf-capture polling for the Welcome line for its full timeout. The wait loop now
uses parity-capture's process check and prints the first Exception/Fatal line. Verified: bash -n,
PacingLogFormatCoverage and TaaAcceptanceHarnessCoverage tests 14/14.
…-floor parity rule, perf-capture OpenGL noise

Both backends ran on the NVIDIA GeForce RTX 4070 Laptop GPU (driver 615.71.09), renderer and GPU
lines confirmed from each client log. Both parity dump paths executed (30 attachments, 53 files).
Pacing medians: OpenGL 6.08 ms mean, 8.51 p99, 0.55 stddev; Vulkan 9.90, 20.08, 4.96, gate fails
on p99, stddev and blocking uploads (the Milestone 1 target). GL-vs-GL launches are not
bit-identical (far shadow map 0.947), so M1.6 is judged against the same session's floor. Found:
SSAO colour1 alpha 1.0 on GL, 0.0 on Vulkan; GetQueryResult flushes the frame about once per frame.
perf-capture no longer prints a missing-stats-file error on OpenGL runs (pinning tests 14/14).
…om the far point

The sky motion pass and the resolve fallback used the unprojected far-plane point as a
direction. With the eye offset inside CameraMatrixOrigin that point is not a pure direction
and biased the sky motion by about 0.56 px at 1.7 m eye height. Both consumers now form
farH.xyz * nearH.w - nearH.xyz * farH.w with a homogeneous sign guard.

Verified: GPU tests (TaaSkyMotionTests still-camera cases at eye heights 1.7 m and 25 m and
pitches 0, -0.35 and 0.2 write at most 0.3 px of sky motion; TaaResolveTests keeps the sky put
under camera translation and above the origin) and the Optimum.Tests coverage on both shader
files. Not yet judged by eye in game.
@NightHammer1000
NightHammer1000 marked this pull request as ready for review September 25, 2026 08:34
@NightHammer1000

Copy link
Copy Markdown
Author

A few Screenshots:

Vanilla AO

Screenshot 2026-09-25 135710 Screenshot 2026-09-25 140414

Custom GTAO

Screenshot 2026-09-25 135728 grafik

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant