A renderer with no binding model
Textures are integers and buffers are addresses. No descriptor sets, no vertex input, no render graph, and no rasterization path beside mesh shaders. How that works
A renderer with no binding model
Textures are integers and buffers are addresses. No descriptor sets, no vertex input, no render graph, and no rasterization path beside mesh shaders. How that works
Data-driven world
A sparse-set entity-component world with a reflection system that powers the editor, serialization, and networking.
C# scripting
Gameplay in C#, with in-editor hot reload, a typed component API, a message bus, and editor tooling built around it.
Editor first
A full editor, the scene outliner, inspector, asset browser, and live tools, on the same runtime as the game.
Most graphics abstractions are shaped by hardware that stopped existing years ago. Binding slots, vertex input layouts, and per-draw descriptor sets are all concessions to fixed-function stages modern GPUs no longer have. Lumina skips the concession and targets the hardware as it is.
Descriptor indexing, buffer device address, and mesh shaders are hard requirements on Vulkan 1.4. Device selection checks for them up front and says so in a dialog when a GPU comes up short. There is one way to draw geometry, one way to reach a texture, and one way to reach a buffer, with no fallback path beside any of them.
A texture is an integer
There are no descriptor sets in the API at all. Sampled images, storage
images, and samplers live in one process-wide heap and are addressed by a
uint32 slot. The heap is bound once per command list, so that integer can
sit inside a GPU struct and be read by a shader nothing ever told about the
texture.
A buffer is an address
There is no buffer object either. An allocation is a 64-bit device address, and a shader chases it the way C++ does, so a GPU struct can point at another buffer. Vertex input went with it. A mesh shader pulls its own vertices through a pointer.
A draw is a pointer
Every draw and dispatch takes one GPUPtr that becomes the shader’s push
constant. Per-draw data is a struct written into a per-frame bump allocator
and handed over by address, which is why moving a decision from the CPU to
the GPU never changes a signature.
Mesh shaders are the pipeline
There is no vertex and index path to fall back to, and no fast path to opt into. A meshlet is 64 vertices and 64 triangles with quantized positions, and meshlet culling is a compute pass that writes its own indirect arguments.
No render graph
Barriers name stages, not resources, so there is nothing to track. No
last-writer table, no aliasing analysis, no per-frame graph cost. Every image
stays in GENERAL for its whole life. Passes run in the order they are
written, so a capture and a profiler zone line up with the call site.
One dispatch per material
Opaque geometry writes a visibility buffer, a compute pass sorts every pixel into its material’s run, and each material shades exactly the pixels it owns with one indirect dispatch. The cost scales with the number of visible shaders, and an instance count of ten thousand costs the same as ten.
const uint32 AlbedoSlot = RHI::HeapWriteTexture(RHI::GetGlobalHeap(), Albedo);
struct FArgs { RHI::GPUPtr Instances; uint32 AlbedoSlot; uint32 Count; };const RHI::GPUPtr ArgsPtr = RHI::CopyTransient(FArgs{ Instances, AlbedoSlot, Count });
RHI::CmdSetTextureHeap(CL, RHI::GetGlobalHeap());RHI::CmdSetPipeline(CL, Pipeline);RHI::CmdDispatch(CL, ArgsPtr, Groups, 1, 1);That is all of it. The heap is bound once, ArgsPtr carries everything else, and
the shader reads Instances as a pointer and AlbedoSlot as an index into the
heap. Nothing is rebound per draw, because there is nothing left to rebind.
The scene is retained, and nothing walks it per frame. Culling is a GPU-resident chain that sizes its own dispatches and writes its own indirect arguments, in two phases against a persistent per-instance visibility set. An early pass replays what the camera could see at the end of last frame and produces the depth a Hi-Z pyramid is built from. A late pass then tests everything else against that fresh pyramid and emits only what the early pass did not draw.
Counts stay on the GPU from end to end. Pixel classification prefix-sums its own material runs, the mesh stage reads a packed draw word the cull wrote, and no stage reports a number back for the CPU to act on. There is no readback, and no frame of latency spent waiting for one.