Birdoggydog's Builds

Measuring a Godot scene against a Steam Deck budget

One bench scene that hides each type of drawn thing in turn, two benchmark ratios that turn desktop times into Deck times, and a budget table at 60 and 30.

Blobber has a frame budget now. It’s written for the Steam Deck at 1280 by 800, with a row for each type of thing the game draws, at 60 and at 30 frames a second. One scene measures everything that gets drawn outdoors at once. Nothing was actually run on a Deck. Every Deck figure is the time measured on my machine multiplied by a ratio. With scatter, rocks, the HUD and 20 dressed orcs, the Deck’s CPU has room at 60 and its GPU doesn’t. At 30 everything fits. I had agents build the bench and write the report.

The measuring scene at its four stations with the crowd of dressed orcs stepping from 0 to 40, a live panel of draw calls and triangles, the Deck’s bars against the 60 and 30 lines, and the budget table.

Before 10 October I didn’t have a goal written down. There was no target machine and no target frame rate. There was a budget for scatter only, built on an assumed machine “8 times slower” than mine, and no scene had ever measured it all together. I think the Steam Deck is a great target. I’m aiming for 60, with a set of targets at 30 to make trade-off decisions against.

Building one scene that draws everything

It’s one Godot 4 scene on a shared rig script. It loads the painted 10 km island exactly the way the exploring scene draws it, with Terrain3D ground, a far-ground mesh and water. It turns the scatter trial on in its fullest mode, with the installed rock models and stand-in meshes for trees, shrubs and huts. It shows the party’s HUD. And it adds a crowd of one type of creature: a bare orc body with eight rigid armor pieces and a sword attached to its sockets at run time.

The crowd’s pieces are taken in rotation from three different gear files. That’s the most expensive way to do it, because fewer pieces are identical. The bodies walk in place. The level’s own monsters are freed so no AI is running. The game clock is stopped at 13:00, the party stands still and the random seed is 1.

The window is 1280 by 800 with MSAA 2x, vertical sync off, and the Forward+ renderer the project ships with. The same run is repeated under the Mobile renderer by starting the engine through a small wrapper that adds --rendering-method mobile.

There are four camera stations. open is a sparse site with the crowd 12 m away. dense is the thickest rainforest the scatter rule makes, with 4.3 million triangles in the frame. across is a ridge looking over the whole island. close is a cove with the crowd 4.5 m away. Each station steps the crowd through 0, 5, 10, 20 and 40 bodies.

How do I find what one class of thing costs?

Everything that gets drawn is filed under one of eight classes: terrain, scatter, rocks, site, bodies, pieces, held (the swords) and ui. At crowds of 0 and 20, the bench hides each class in turn and measures again. A class’s share is the whole frame minus the frame with that class hidden. Whatever’s left with every class hidden is the empty frame: sky, tone-mapping, the MSAA resolve and SSAO.

This doesn’t need any instrumentation inside the renderer. The downside is that a time share is the difference between two noisy numbers. The time shares can miss the whole frame by up to 1.2 ms, because the CPU and the GPU work at the same time. The shares of draw calls and triangles do add up, to within 0.8% and 0.4%.

Each row is 240 frames after 40 frames of warm-up. It records frame time (mean, 95th percentile and worst), the renderer’s time on the CPU and on the GPU, draw calls and triangles in view, shadow passes, object count and video memory. The two render times come from the rendering server’s viewport render-time measurement, switched on for the root viewport. The engine’s Performance monitors for process and physics time only refresh once a second, so I dropped them.

One run is 98 rows and takes about 100 seconds unattended. The rows are written out as JSON. A Python report takes the median of three runs and prints every table in the budget document.

Turning desktop milliseconds into Deck milliseconds

I use two ratios. Each one comes from a public benchmark and has a stated range.

  • GPU: 17, and it could be anywhere from 10 to 25. That’s 3DMark Time Spy graphics, 27,920 for my card against about 1,617 for the Deck. It could be as low as 10 because a big card that’s 95% idle doesn’t run at full clock, and every pass has a fixed cost. It could be as high as 25 because the Deck’s memory is shared, and its 15 W is shared with the CPU.
  • CPU: 2.2, and it could be anywhere from 1.8 to 3.0. That’s Geekbench 6 single core, about 2,500 against about 1,300, made a little worse to allow for the shared power and a different Vulkan driver.

The empty frame, the terrain shader and all triangle cost are scaled by the GPU ratio. Every draw call, the HUD, skinning and scripts are scaled by the CPU ratio. Draw calls, triangles and megabytes are counted, not scaled.

The old “8 times slower” assumption was wrong in both directions. It was 3.6 times too harsh on draw calls and half as harsh as it should have been on triangles. So its conclusion that draw calls run out first is backwards for the Deck. Triangles and shader cost run out first.

Where does the frame stand at 60?

I keep a fifth of each frame spare, so the CPU and the GPU each get 13.3 ms at 60 and 26.7 ms at 30. They work at the same time, so each one gets the whole frame, and the frame takes as long as the slower of the two.

Worked out for the Deck with 20 orcs, the CPU takes 7.4 to 7.9 ms and the GPU takes 14.4 to 20.0 ms. These are the GPU’s rows at 60:

RowBudget (ms)Today (ms)
Empty frame: sky, post, MSAA3.04.6 to 5.2
Terrain and far view4.05.5 to 6.3
Scatter and flora2.50 to 5.5
Rocks0.50 to 0.6
Creatures, a crowd of 201.50.5 to 3.2
HUD1.01.3 to 1.9

The terrain only has half a million triangles, so its cost is the shader. Scatter in a dense stand is 3.7 million triangles against a budget row of 1.3 million, and half of those are shadow passes. The HUD is the biggest single source of draw calls, 374 against a budget of 150. That’s a third of the frame’s calls with nothing happening on it.

Turning the sun’s shadows off gives back 1.9 to 3.4 ms, a third of the draw calls and half the triangles. MSAA 2x costs 1.0 to 1.9 ms and 77 MB.

The Mobile renderer runs the scene unchanged with one warning (no SSAO). It’s a fifth to a quarter cheaper on the GPU, at 11.1 to 15.1 ms, and uses 90 MB less. But it draws every worn piece as a separate call. 20 dressed orcs are 200 calls in view on Mobile and 73 on Forward+, because Forward+ batches identical pieces. I haven’t judged how it looks yet, and I’m leaving the choice open for now.

Does a worn piece need 150 triangles or 240?

It doesn’t matter which. 160 pieces at 240 triangles is 38,400, and at 150 it’s 24,000. Each one gets drawn about 2.5 times with the shadow passes, so the difference is 36,000 triangles in the frame. That’s 0.06 ms on the Deck’s GPU.

What costs is how many different pieces are in view. It’s one draw call for each different piece in each pass, no matter how many bodies are wearing it. One suit on 20 orcs is about 50 calls. Pieces from three files are 92 to 167 calls, which is 0.2 to 0.8 ms on the Deck’s CPU. The body is 3,576 triangles, which is two thirds of a dressed orc.

I kept the low target anyway, for a different reason. My eventual dream is a lot of big packs on the screen.

Next

I haven’t measured a real Deck, interiors, buildings in view, monsters with their AI running, or video memory by class. There’s also a frame of 32 to 38 ms that shows up in the second after the first five dressed orcs arrive at a station, and I don’t know why yet.