Frankincense wrote: Sat Sep 26, 2020 5:37 pm
Current HLMS implementations use a large buffer of data with an index to perform the automatic instancing, should a const buffer always be used? If not, when should it not be used?
This varies per GPU, but the general rule is that const buffers are better if the index (e.g. the value of 'i' in data[ i]) is the same for a lot (best case: all) of the vertices or pixels while the other buffers are better if the index is highly dynamic (e.g. 'i' can be different for every vertex/pixel)
Also if you need far more than 64kb, then const buffers are not an option.
You can read more from my post in my personal website in
UBO vs TBO (UBO = ConstBufferPacked, TBO = TexBufferPacked)
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]What are the shader buffer commands (CbShaderBuffer) and do they relate to anything in the old material system?
They just bind from C++ the buffer to a specific slot so that it can be available for shaders using that slot.
The type of buffer the shader expects must match what C++ sets (e.g. if shader expects a texture buffer at slot 3, then C++ must set a TexBufferPacked at slot 3)
No, there is no equivalent in the old material system
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]I assume 'setProperty' would be equivalent to setting define flags when building a shader?
Yes.
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]How are textures 'baked' and what purpose does this serve?
We try to group together as many textures as possible. We use texture2DArrays for that purpose because they're widely supported in HW. The main restriction is that textures that can be grouped together must have the same format and resolution.
The reason for grouping is performance. Rather than doing:
Code: Select all
for each object
{
setTextures( textures[i], num_textures[i] );
draw( object[i].vertexCount );
}
We perform:
Code: Select all
for each group
{
setTextures( group[i].texturesArray );
for each object in group
draw( group[i].object[j].vertexCount );
}
This way we call setTextures() far less frequently (the call instruction isn't that expensive; but there is a lot of work that needs to be done on CPU and GPU sides when swapping too frequently).
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]Does each datablock generate a hash so there is only one copy if all of the data is the same?
Yes-ish... and no.
This is explained
in detail the manual:
There are two components that needs to be evaluated that may affect the shader itself and would need to be recompiled.
- The Datablock/Material. Does it have Normal maps? Then include code to sample the normal map and affect the lighting calculations. Does it have a diffuse map? If not, avoid sampling the diffuse map and multiplying it against the diffuse colour, etc.
- The Mesh. Is it skeletally animated? Then include skeletal animation code. How many blend weights? Modify the skeletal animation code appropiately. It doesn't have tangents? Then skip the normal map defined in the material. And so on.
When calling Renderable::setDatablock(), what happens is that Hlms::calculateHashFor will get called and this function evaluates both the mesh and datablock compatibility. If they're incompatible (i.e. the Datablock or the Hlms implementation requires the mesh to have certain feature. e.g. the Datablock needs 2 UV sets bu the mesh only has one set of UVs) it throws.
If they're compatible, all the variables (aka properties) and pieces are generated and cached in a structure (mRenderableCache) with a hash key to this cache entry. If a different pair of datablock-mesh ends up having the same properties and pieces, they will get the same hash (and share the same shader).
In short if two datablocks are identical AND the mesh properties are identical (e.g. vertex format, animation), then the hash will be the same.
This also means that the same datablock applied to two different meshes may have the same hash, or different hashes.
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]How do the 'calculateHash' functions actually calculate the hash? It seems that they mostly just set properties in the shaders?
The same section of the manual explains it, but it basically analyzes data required (e.g. does it use normal maps?) analyzes the mesh structure (does it use float3 for position? or does it use half4? does it have normals?) and sets the properties.
The hash is calculated on properties set.
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]What should be inside 'calculateHashForPreCreate' vs 'calculateHashForPreCaster'?
Whatever the derived implementation needs to set. For example HlmsPbs checks if the material needs normal mapping there and if so, sets the properties; then the shader uses this property to add shader code to deal with normal mapping.
HlmsPbs also checks one by one which textures are needed.
The HlmsUnlit implementation doesn't have normal mapping so it does not have such code.
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]What does 'PreCreate' and 'PreCaster' mean in this context?
Each mesh-material combo needs mostly (at least) two shaders: the main one used for rendering; and the one used during shadow casting.
The shadow casting shader tends to be extremely simple, but it may need to still account for skeletal animation and if alpha tested shadows are required (e.g. foliage shadows) then the caster needs to account some textures to read the alpha.
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]Why are some properties set in both 'calculateHashForPreCreate' and 'calculateHashForPreCaster'? How should I decided which one to put my properties in?
This was answered by the previous question: former is what's needed for main rendering, latter is only what's necessary for shadow casting.
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]Are these functions called only once per object being rendered, or once per unique datablock on creation?
When the material is assigned to the Object. Changing a material to the object casues this function to be called again (which may end up reusing an existing hash in the cache, or creating a new one).
The idea is to move most work to creation time (although there is a final step that inevitably happens at runtime every frame, per object: Hlms::getMaterial gets called and if the two hashes haven't been merged into the final hash, then a new shader is generated; else a shader in a cache is provided)
Frankincense wrote: Sat Sep 26, 2020 5:37 pm
[*]In 'HlmsPbs::calculateHasForPreCaster' a number of properties are removed from mSetProperties, how are these chosen and what's the effect?
Because the caster version is usually a heavily watered-down version; it's easier to remove unnecessary properties rather than evaluate them again.
Technically we could leave all properties set and add an additional that says "is_caster = true" (which we do, btw) so that the template follows a different code path.
However this prevents shader reuse. Two material-object combos may end up with exactly the same shader; however if their properties set are different (even if they're ignored) they'll get different hashes, and if they get different hashes, they'll be treated as different shaders (which hurts performance). That's why we prefer stripping down properties as much as possible
[*]What is done in 'createShaderCacheEntry' and what should go in it vs the 'calculateHashX' functions?
calculateHashX gets called while the shader hasn't even been generated. And properties set this way will be part of the hash.
The function createShaderCacheEntry creates the actual shader; and setting properties here won't affect the hash; which is usually a bad thing (because two different shaders will be seen as the same shader).
[*]When is 'preparePassHash' called? I assume its setting all of the data to be passed to the shader each time a shader pass is rendered (e.g. main pass + shadow caster)?
Also
explained in the manual, preparePassHash gets called per pass to set properties that are global to all mesh and materials for that pass.
This means there are 3 hashes:
- Material-mesh pair combo
- Pass
- The final hash, which combines the former two
[*]Does this affectively replace the old 'automatic' material properties to pass this data?
I don't know what are the "old 'automatic' material properties"
[*]In the Datablock 'setter' functions, when should I call 'scheduleConstBufferUpdate' vs 'flushRenderables'?
flushRenderables: Whenever the shader must change. For example the material uses diffuse textures, and now you remove them. This requires the shader to change (a property needs to change its value, or be set/unset). Hence flushRenderables will force all objects using this material to calculate their hashes again (it's basically like unsetting the datablock and setting it again for every object to force a rebuild).
scheduleConstBufferUpdate: when material data living in GPU memory needs to change; without requiring the shader to be recompiled. For example if you set the diffuse colour from blue (0, 0, 1) to red (1, 0, 0); the GPU memory needs to be altered with this change.
It's essentially memcpy( gpu_memory, cpu_memory ); with the new parameters.
[*]Are the macro/sampler blocks part of the hash generated?
Macro & Blendblocks: Yes. The reason is more technical, has to do with how PSOs in modern APIs (Vulkan, Metal & D3D12) work (see we use mLifetimeId).
Sometimes the reason is less technical. For example Hlms checks the blendblock->isAutoTransparent() to see if the shader should be aware that alpha blending is required.
Samplerblocks: Yes, but only the amount of sampleblocks matters (because it results in a different shader)