Parity
The Genshin area recreates the game's own screens, so "does it look right" has an exact answer: the game. Parity is the loop that asks it. It is built for an agent, which reads images but drives no browser of its own: every step is one command that prints numbers and writes an image the agent reads. The user's eye stays the final check. The loop is fast enough that nobody has to wait on that check.
How it works
flowchart TD
WIKI[Wiki images and animations] --> FETCH[fetch / frames]
GAME[The installed game: the user plays, the tool records] --> CAP[still / record]
CAP --> FRAMES[frames: fixed-rate stills and a contact sheet]
FETCH --> REF[(References, outside the repository)]
FRAMES --> REF
REF --> TRACE[trace: a mark as one SVG path, full resolution]
TRACE --> SCREEN[Screen component and its fixture, in the world package]
REF --> MEASURE[measure / zoom: colours, sizes, positions]
MEASURE --> SCREEN
SCREEN --> PAGE[Parity page on Vite, no Nuxt]
PAGE --> COMPARE{compare: reference, ours, difference, and a score per cell}
REF --> COMPARE
COMPARE -->|off| SCREEN
COMPARE -->|matches| VISUAL[Visual suite: our own image, committed beside the component]
- References stay outside the repository. They are HoYoverse's images, so they are fetched and recorded into
~/Esposter/genshin-parityand never committed. The repository holds only what each screen is judged against (ParityReferenceMap): a wiki file, or one frame of a recording, named by the recording and the second it is taken at. - Search before recording. Most of what a screen needs is already published: the wiki holds the game's screens and its standalone marks, the community's data dumps hold its in-game text, and public videos show the rest, including the English client when the recording is in another language. A clip from one is fetched into the captures folder with yt-dlp from a throwaway virtual environment, and sampled like any recording. A recording of the installed game is for what nothing published shows, such as exact timings at 60 frames a second and colours read true.
- The installed game is recorded, never driven. The user starts the game and plays while the tool records. It sends the game no input, since the game's terms count injected input as botting.
- Only the game's window is recorded. Capture goes through Windows Graphics Capture, found by the game's executable. Nothing else reaches the file: no window above it, no desktop or webcam behind it, no cursor, and no overlay that is its own window. A recording of the screen's rectangle was tried first and took whatever was on screen whenever the user switched away.
recordwaits for the window to open, so a recording can be armed before the game starts and take it from its first frame. It encodes on the GPU at 60 frames a second, near-lossless, into Matroska, which stays readable if the recording is cut short. A minimised game gives no frames, so a recording also ends by the clock, at its length in real time, rather than waiting for frames that are not coming. - FFmpeg is pinned, not installed. Window capture arrived in FFmpeg 8, which no npm package bundles. The tool fetches one versioned release the first time it needs FFmpeg, refuses it unless its SHA-256 matches the pinned one, and unpacks it into the scripts package's
node_modules/.cache, so no install or CI run ever pays for it. - A vector is used as it is, and only a mark with none is traced. Wikimedia Commons holds some of the game's logos as SVG, the publisher's among them, and its paths are drawn directly, so no edge can chip.
- A glyph is traced from the game's own mark. The tracer starts from the wiki's standalone render of a mark where one exists, since a mark on a screenshot is small and, when pale, too faint to split cleanly. It traces at the source's full resolution, never reduced, and prints the region beside the trace to check. Ink is told from paper by how far each pixel strays from the colour at the region's corners, so a mark dark on white, lit on black or coloured on either traces alike, and a run of ink or a hole in it smaller than a sliver of the region, a sparkle printed on a logo, a speck of compression or a pale sparkle punching through a letter, is cleared from the split before it is traced; every path the trace finds is then kept, so no small part of the mark is lost. The path it writes is the one the component draws. Nothing from the game ships except shapes derived this way, a choice made so the screens match exactly.
- A screen is a component folder with a fixture, in its section. Every screen is a presentational component in the world package, laid out as Nuxt lays out the app's:
components/<Section>/<Name>/Index.vue, named by its path (Login/ScreenisLoginScreen), with itsIndex.fixture.ts, its browser tests and its approved images beside it. The sections follow the game (Game,Splash,Loading,Login,World), andGame/Openingshows the game's first screens in turn and owns every handoff between them. Having a fixture is what puts a screen on the parity page and in the visual suite, named as the package's own barrel exports it, so there is no list to keep. A fixture'svariantsare further states it is approved in, such as each stage of the login's interface. A fixture markedisMotionOnly, such as a 3D scene whose anti-aliasing jitters every frame, is shot on the page but kept out of the suite. A screen draws insidegenshin-interface'sGameScreen, whose--unitscales the game's 1920 by 1080 screen by the window's smaller axis, as the game scales its interface; scaling by height alone was tried first, and a window narrower than 16:9 set text sized for the height in a box sized for the width. A screen also draws no heading element, since a host page's heading styles (the console's gold, its inherited font) reach into it; a heading is a paragraph withrole="heading". The pieces screens share aregenshin-interface's, held by a suite of their own. - The comparison is numbers first.
compareshoots the screen in the machine's own Edge at 1080 CSS pixels high, scaled to the reference's pixel size. It prints the mean difference and a six-by-six grid of per-cell differences, then writes reference, ours and their difference side by side. A cell that stands out says where to zoom. It prints two scores beside them for a scene, which is rebuilt from shapes rather than copied and so never matches pixel for pixel: the shape, as the share of edges each image has where the other has one too, and the tone, as the difference of their colours blurred past any texture. A reference can hand its screen props of its own, so one screen is judged in each state a reference shows, such as a scene at each time of day. - The frame is calibrated by regression, never searched.
calibratefits the frame-wide terms over the pixels the witness draws its parts on, in the order light meets the eye, each held for the next: the sun's and the ambient light's colours by least squares from the albedo and the normal (linear in them for one sun, whose heading and elevation are the one outer solve, a grid then the simplex, kept between 5° and 85° of elevation since grazing light lets a vast sun explain a sliver of pixels), the fog's colour and density from the depth (held clear where the drawn depths barely spread, since one depth cannot tell density from colour), then each channel's grade as a monotone curve from the predicted colour to the reference's, made monotone by pooling adjacent violators, beside plain sRGB's residual. One pass, never iterated: a grade fitted over a poor light regresses toward a flat curve, and reading the next pass's reference through it blacks it out. Given the game's grading tables, each is scored against the predicted colour. It reads a solved pose, so a pose that lands one part alone calibrates that part's light alone. - FLIP is the approval number. NVIDIA's standard dynamic range FLIP is ported as it stands (
scoreFlip): contrast-sensitivity filtering in YCxCz at the viewing distance's pixels per degree, the HyAB distance in Hunt-adjusted CIELab, and the luminance's edges and points, each pixel's error the colour difference raised to one less the feature difference. Its test holds it to NVIDIA's own evaluator, and on FLIP's example pair it reads 0.1597, as its readme does.comparereports it beside the mean, read at the structure's width as on a full screen. With--witness,compareandattributealso score each layer apart (scoreLayers): a family's pixels in the witness's part target, and the sky where no part is, each with its own shape, tone, detail and mean FLIP, so a change is judged where it lands.attribute's committed loss table is each stand-in's FLIP loss on each layer against the witness's. - The scores are committed, as a bench's are. Every
comparerewrites its reference's row ofParityReferenceMap.snapshot.md, beside the map, andcompare --allrewrites them all, so a change's diff shows what it moved on every screen, and the map's test holds every reference to a screen the page shoots. - Motion is matched frame by frame. A recording is first sampled at one frame a second to find its events, then each event is read at its full rate over its own window of time. Our screen's motion is held at its start on the parity page and shot at the same moments.
entryholds the screen's own animations as it mounts.propslets those finish, then applies the fixture'smotionPropsand holds the transitions they start; a still and the visual suite show the fixture's first state.lumaprints one region's darkness across both sets of frames as two curves, so a fade's duration and a wipe's steps are read off numbers rather than raced. Window capture takes a frame only when the window changes, so a curve from it jitters by a frame or two where frames were skipped, and a fit is judged over the whole curve. A scene's own motion runs on its frames' clock rather than on animations, sofilmfakes the page's clock and moves it a frame at a time, shooting the screen at exact moments however slowly the page draws, with props set at moments to play its stages as a click or a load would, into stills and a contact sheet beside a recording's sampled at the same rate. With--witnessit films the exports in place of our parts. - A scene's cost is benched, not guessed.
benchdraws the screen as fast as the GPU allows rather than at the display's refresh, and prints its frame time's median and slow tenth against the main thread's busy time a frame, which tells a scene held up on the CPU from one held up on the GPU, then the renderer's passes, draw calls and triangles a frame, the objects it draws by kind, and the pipelines, geometries and textures it keeps, which grow from one bench to the next in a scene that leaks them. The page reads them from the renderer the scene hands any host that asks (SceneContextKey). The login's hundreds of cloud billboards and its walkway's pieces as meshes of their own were found this way, and drawing each band as one instanced sprite and the walkway as one batch took its frame from 2.8 to 1.6 milliseconds. - Nothing is approved uncompared. The visual suite only holds a screen to its own last image, so it says nothing about the game. What does is
compareagainst a reference in the same language and build, and a test fails for any screen with a fixture that no reference names. A title logo sized from another language's client was approved once without a reference and came out nearly twice the size of the English client's, which is why the test exists. A public video that letterboxes the game is cropped to its screen by the reference'scrop. - An approved screen is held by the visual suite.
pnpm test:visualrenders every screen and every variant from its fixture in Vitest's browser mode and matches it against its own last image,Index.<platform>.pngandIndex-<variant>.<platform>.pngbeside the component, as a test sits beside its code. It is a separate config, so an ordinary test run never starts a browser, and it reads the engine and the interface library from their source, so a change there shows without a build. The same config runs the tests of what only a browser can play:Game/Opening's plays the whole opening, the login's flight included, and checks every handoff, so a screen skipped along the way fails a run rather than a reader's eye.-uapproves what the screens draw now. A new image's first run writes it and fails once, which is Vitest's way of asking for that approval. - An interface over a scene is compared over the reference's own frame. A reference with
isBackdrophas its frame drawn behind the screen, served to the page by the shooting browser and drawn pixel for pixel, so every cell the interface leaves bare scores exactly 0.00% and only the interface can differ. A capture's frame is rewritten without the gamma and primaries FFmpeg tags it with, since a browser colour-manages a tagged image and drew the frame darker than its own file. Glass the interface draws translucent (the prompt band) lands on the recording's own glass and darkens twice, so its colour is solved from the scene around it, and the score holds its opaque pieces and every placement. The login's interface reached a few tenths of a percent on each stage this way, the rest being the older build's build string and repair button. - A witness render draws the exports through our own scene. With
--witness, the shooting browser serves a component's exports to the page and the scene draws them in place of its own parts (SceneWitnessKey), under its own camera, light and frame, so a stand-in of ours is priced against the game's own data (scene derivation). Nothing of it is bundled: the page only fetches what the shoot serves it.posesolves the camera from landmarks, points of the exports whose places the witness knows (a share of a part's bounding box), against the pixels a reference names for them, each snapped to its nearest corner: from six or more the projection is read in closed form by a direct linear transform, from fewer from a start with the field of view held at one read off two widths, and Levenberg–Marquardt then minimises the reprojection error, printed per landmark. A few steps of the simplex then refine it, the held axes still held, on the distance from the chosen families' silhouettes, taken from the part target rather than the shading, to the reference's edges, so the clouds, which draw no part, cannot pull it. The seams between one family's parts are left out, since the reference draws them as cracks or not at all: the walkway's blocks meet along its painted cracks. Two parts in one frame share one camera, so a part whose own landmarks barely fix an axis takes it from one that does: the door's three points trade its pitch against its height, and the walkway's wings fix it.trackdoes the same at each sampled frame of a reference's recording, each from the frame before, for a path no clip holds.place --landmarksfits a row's offset and turn to landmarks on it through a camera already solved, by least squares on their reprojection with no frame drawn, and names the landmark a pairing went wrong on by its error laid out;--axesholds every axis it does not name, and an edge landmark (a round part's silhouette, its bounding box's side halfway through its depth) counts only across.placewithout them holds the camera and refines where a group of families stands, one offset shared by them, on their boundaries' distance to the reference's edges and the reference's edges' to them, those taken where the other families do not stand: one way alone rewards drawing less, a phase that carries every near part away scoring best. A row the script scrolls has its phase read off the parts the reference shows first:partsnumbers each part where it lands over the reference and the witness's render, andview --offsetstries a phase, the login's lantern tower putting its row 144 metres along at the door, whichplace --startthen refines. From its laid-out place a refinement on edges settles wherever the reference's clouds and haze pull it, as a 0.43-metre lowering did before the phase was read. The scene stands the witness's rows where its own scroll, so a tool sees the exports where the scene stands its parts. With--top-row,poseandtrackprice only the edges below a row, clear of what the exports draw otherwise (the login's walkway assembling itself at its far end). The witness lays out every copy a spawn scrolls, so a camera gliding past a block's end still has the next copy under it. Where the camera holds still and the world glides toward it,glidereads the pace off the recording instead: at the camera's pose every row of level ground is a distance (readGroundRow), so a column of the ground straight ahead, resampled into metres and correlated frame to frame (measureGroundShift), is the metres the world moved, summed over each window into metres a second beside the frames the recording held. The grid search over a line distance this replaced absorbed a wrong arrangement into plausible poses. - The witness settles in one frame and writes a G-buffer. On the parity page the witness's post chain resolves edges with SMAA, which keeps no temporal history, and setting a view holds the scene's clock, so one drawn frame is the view and the same view draws the same frame.
gbufferthen renders the witness's parts alone into floating-point targets (the normal halved and lifted into 0 to 1, since a material's colour output clips what falls below 0, and let down again as it is read) read back from the renderer (renderWitnessTargets): the depth along the view, the world normal, the unlit albedo and each part's identifier with its family, named in the header. A pose's refinement reads the part target alone, and each part's target materials are built once and kept, since a material built afresh each read is a pipeline WebGPU compiles and keeps until the page crashes. The part target is the reference's segmentation at that pose, which the later tools mask by, andoverlaydraws its families' boundaries over the reference beside a map of the reference's edges by how far each is from them, printing each family's distance, so a part that does not land shows where. - The world's shapes are measured off the game's own assets, never copied. The interface is rebuilt from screenshots, which its overlay compare holds exactly; a scene's towers, doors and profiles come from the installed game's files. AnimeStudio exports them with the game closed into the references folder, outside the repository like every other reference, and a transform fits our own kits' parameters to them, so only numbers of ours ship (derived assets).
Deriving a screen
A new screen, or a new piece of one, is derived in one order, so every source is found once, recorded where it is used, and never searched for again:
flowchart TD
P["Search what is published: the wiki, data dumps, public videos"] --> REF["References: ParityReferenceMap, and the component's Index.reference.ts sources"]
P --> B["Find the screen's blocks through an indexed asset: a mesh, a clip, a material"]
B --> MAP["Its DerivedAssetComponentMap entry: names, roots, interface root and anchor, clips"]
MAP --> RUN["genshin:assets extract, shaders, inventory, interface, clips, witness"]
RUN --> REF
RUN --> K{What is the piece?}
K -->|"interface"| UI["Placed from the interface tree's rects, moved by its decoded clips"]
K -->|"3D"| AR{"Width ratios of neighbouring parts match the reference?"}
AR -->|no| PAR["Fix the arrangement: a lost parent's place and scale"]
PAR --> AR
AR -->|yes| POSE["Read the pose by perspective from parts of known size"]
POSE --> W{"Witness render at that pose lines up?"}
W -->|no| AR
W -->|yes| SD["The scene derivation: calibration, loss table"]
UI --> C["compare against its references"]
SD --> C
C -->|off| K
C -->|matches| V["Approve in the visual suite"]
V --> R["Record: findings in the reference, conventions in the skill"]
- Read before searching. A component's
Index.reference.tsand the game data formats page's shortcuts come first; a search whose answer is recorded is not run again. - Every derived value cites its source. A rect, a curve, a fitted shape or a constant names the key of the reference source it is taken from.
- Exact data outranks measurement. A RectTransform's anchor, a clip's curve or a shader's program is used before a position, a timing or a model measured off a recording; a recording measures only what is fieldless, such as a layout group's spacing or a script's settings.
Commands
From scripts/, as pnpm genshin:parity <command>, each with its own --help; the parity page from packages/genshin-world, as pnpm parity.
| Command | What it does |
|---|---|
fetch | Every reference not yet held, as PNG |
compare <reference>, compare --all | Shoots its screen, prints the scores, writes reference, ours, difference, and its report row |
compare <reference> --witness <component> | The same over the witness render, its image kept apart and its scores out of the report |
gbuffer <reference> --witness <component> [--pose] | The witness's depth, world normal, unlit albedo and part per pixel in one settled frame, as raw floats with a header and a preview |
overlay <reference> --witness <component> [--pose] | The witness's family boundaries over the reference, the reference's edges by their distance from them, and each family's distance |
pose <reference> --witness <component> | The camera pose from the reference's landmarks, its reprojection error per landmark, then optionally refined on families' edges |
track <reference> <time> <seconds> --witness <component> | The pose at each sampled frame of the reference's recording, refined on families' edges from the frame before's |
shoot <screen> <w> <h> [--motion entry] [ms…] | The parity page's screen at a size, paused at each time when given |
film <screen> <seconds> [--at json] [--fps] [--witness] [--beside rec@s] | The screen on a faked clock at exact moments, props set at moments, each still over the recording's frame then when given |
bench <screen> [w] [h] [--props json] [--frames] | The scene's frame time and main thread's share, its draw calls and objects by kind, and the resources it keeps |
frames <file or File:…> [fps] [start] [seconds] | A video or animated image as stills and a contact sheet, over a window |
luma <x> <y> <w> <h> <image or folder>… | One region's darkness across images, or a folder's frames, as a curve |
measure <image> <x,y>… | The image's size and the colour under each point |
zoom <image> <x> <y> <w> <h> [scale] [--with images] | A region enlarged with hard edges, the same region of each other image stacked under it |
light <reference> --witness <c> | Each colour channel's shares of the sun's and the sky light's strengths, from the near parts' upward faces and their upright ones turned from the sun under each light alone, beside the sun's direction regressed on the faces' shading and the grid's median residual that says whether it is settled |
rank <reference> --witness <c> | Every term of the reference's error ranked by its ceiling, the most of the frame's FLIP it could recover (its pixels' error summed over the frame's): each family's stand-in, its error over the exports' on its pixels, apart from the light, haze and grade the exports carry too, near, middle and far and by how its faces turn to the light, and the sky by thirds of the frame, each third split into the clouds the reference shows over ours and its clear sky; compare --witness prints each layer's ceiling the same way. A second table draws each stand-in beside the exports at the same camera, moment and light and scores ours against them per family, the stand-in's own shortfall, which a soft recording hides: its FLIP, and its multi-scale structural similarity (scoreLabelSimilarity, 1 identical), which charges detail a few pixels off once at its coarser scales where FLIP charges it twice, so a carved surface aligned but imperfect reads closer than a bare one; the two frames are written one over the other (<reference>.stand-in.png) for the eye. The next term worked is the largest |
plan <reference> --witness <c> --family --least --size [--resolution] | A family's unlit albedo from straight above over a rectangle of the ground, so many pixels a metre, its columns toward +x and its rows toward +z, for a surface's design read in metres; the rectangle is lengthened to a whole multiple of the page's height and widened to one of the readback's sixteen-pixel rows |
cover <reference> --witness <c> | The share of each cloud band the scene draws at the reference's hour, solved by the simplex on the sky's cover band by band of its height over the horizon, ours read as the reference's are: a cloud standing elsewhere than the reference's costs nothing here, where a score comparing pixels charges it twice |
clouds <reference> --witness <c> | The sky's clouds' lit and shaded colours where ours and the reference's stand in different places: ours drawn with the clouds black, then shaded white, then lit white, so each pixel's sky behind and its shares of the two colours are read apart, the reference's clouds those standing brighter than ours without clouds, and the two sets matched quantile by quantile, ours ordered by how lit they are and theirs by brightness; then both skies' clouds by their statistics, each against its own clear sky, a smooth surface fitted under its clouds by least squares that weighs a pixel over it a tenth as much as one under it and a cloud not at all (fitClearSky): their cover, brightness, edge sharpness and spread, and their cover band by band of their height over the horizon, with both frames and their clouds written for the eye (<reference>.clouds.png) |
fog <reference> --witness <c> [--light] | The haze's density and its own and sunward colours, and with --light each light's share per channel from ours under the sun alone and the sky alone, the bins split by how their faces turn to the light: the parts' pixels past its start, the reference's and ours without fog taken back through the tone mapping, binned by depth and by angle to the sun and read by their medians, the colours a linear solve at each density and the density refined by golden section on its residual, toward the fog's own direction and the sky's sun. Where our parts stand far darker than the reference's without fog, it runs the density up to swap them for fog, so the light is settled first |
haze <reference> --witness <c> | Each depth band's median colour over the parts' interior pixels, the reference's beside the witness's drawn under our light without its fog and with it |
exposure <reference> --witness <c> | How much brighter the reference's parts stand than ours, the median linear luminance over the parts' pixels |
sky <reference> --witness <c> | The reference's sky as the game's sky shader draws it, its colours by least squares, its shape refined, beside the reference |
view <screen> --camera x,y,z,yaw,pitch,fov --witness <c> | The scene from any camera, our parts beside the exports they stand for |
place <reference> --families <f,…> --witness <c> | Where a group of families stands on the reference, the camera held: one offset refined on their edges both ways, from --start |
parts <reference> --family <f> --witness <c> | Each part of a family numbered where it lands, over the reference and the witness's render, for landmarks and phases |
polar <image> <bands> <angles> | A ring mark's colours about the image's centre, by radius and angle |
trace <image or File:…> <x> <y> <w> <h> [scale] | A glyph as one SVG path, and the region beside it to check |
launch, still, record <name> [seconds] | Start the game, capture its window once, or record it once it opens (two minutes unless told) |
What it costs
Each step takes seconds, so the loop runs as often as a test would:
- The parity page is up in about a second.
- A comparison, shot included, takes a few seconds.
- The visual suite's first screen takes a few seconds, and each further screen a fraction of one.
- A trace at a mark's full 1600-unit resolution takes a second or two.
The startup loading screen reached a mean difference of a few hundredths of a percent in two edits. The remainder is antialiasing along the edge of the lit marks.
Key files
| File | Role |
|---|---|
scripts/src/services/genshinParity/commands/genshinParityCommand.ts | The commands, one file each beside it |
scripts/src/services/genshinParity/ParityReferenceMap.ts | Each reference, the wiki file it comes from and its screen |
scripts/src/services/genshinParity/compareScreen.ts | Shoot, score and lay reference, ours and difference side by side |
scripts/src/services/genshinParity/shootScreen.ts | The parity page's screen in Edge, over a reference's own frame when it is a backdrop |
scripts/src/services/genshinParity/traceImage.ts | A mark traced into one path at full resolution |
scripts/src/services/genshinParity/solveCameraPose.ts | The camera pose from correspondences: the direct linear transform, then Levenberg–Marquardt |
packages/genshin-world/parity/witness/loadWitness.ts | The exports laid out as the scene's parts, on the parity page |
packages/genshin-world/parity/screens.ts | Every screen with a fixture, for the page and the suite |
packages/genshin-world/parity/screens.visual.ts | The visual suite |
packages/genshin-world/src/components/Loading/Startup/Index.vue | The game's startup screen: the seven marks, wiped in as loading goes |
packages/genshin-world/src/components/Splash/Sequence/Index.vue | The game's splashes: its logos and health notice, timed from a recording |
packages/genshin-world/src/components/Game/Opening/Index.vue | The game's opening: the splashes, the login screen, then the startup screen, handed on by their own timings |
Sources
- Visual regression testing, Vitest:
toMatchScreenshotand where its reference images are kept. - imagetracerjs: the public-domain tracer behind
trace. - Photo Mode, Genshin Impact Wiki: capturing the game without its interface, for the world's references.