[2.1] Performance Optimizations for VR
-
xrgo
- OGRE Expert User

- Posts: 1148
- Joined: Sat Jul 06, 2013 10:59 pm
- Location: Chile
- x 169
[2.1] Performance Optimizations for VR
Hello!!
I currently working on a project that uses Oculus Rift for VR and Ogre 2.1.
I am doing distortions on my own using shaders, not the Oculus SDK, and performance is pretty good with Ogre 2.1, but I really think it could be better yet! I read the blog post of the new unity 5 that mention some features for VR (http://blogs.unity3d.com/2015/06/09/unity-5-1-is-here/). So I have a few questions:
1. (this one is not necessary vr related, but anyways...) I have a warehouse in my scene, in the inside there are a lot of racks, if the camera is outside the warehouse and I look in the direction of the warehouse my fps goes down a lot, because of the racks (lots of polys), but the racks are behind a wall. Is there a way to avoid rendering those objects? some culling method? z-prepass? (or the new unity 5.1 "single pass culling")
2. Shadows! this feature is expensive, but really necessary. In my project the bottleneck seems to be in the polycount (since I tried change the pixel format, pcf quality, and resolution and I don't get much increase). Taking that in count is there a way to reuse shadow maps? (like the new unity 5.1 "shared shadowmaps") What about "SHADOW_NODE_REUSE" does this helps with VR? how do I enable this feature in my shadow node?
3. Any other tips??
Thank you so much!!
I currently working on a project that uses Oculus Rift for VR and Ogre 2.1.
I am doing distortions on my own using shaders, not the Oculus SDK, and performance is pretty good with Ogre 2.1, but I really think it could be better yet! I read the blog post of the new unity 5 that mention some features for VR (http://blogs.unity3d.com/2015/06/09/unity-5-1-is-here/). So I have a few questions:
1. (this one is not necessary vr related, but anyways...) I have a warehouse in my scene, in the inside there are a lot of racks, if the camera is outside the warehouse and I look in the direction of the warehouse my fps goes down a lot, because of the racks (lots of polys), but the racks are behind a wall. Is there a way to avoid rendering those objects? some culling method? z-prepass? (or the new unity 5.1 "single pass culling")
2. Shadows! this feature is expensive, but really necessary. In my project the bottleneck seems to be in the polycount (since I tried change the pixel format, pcf quality, and resolution and I don't get much increase). Taking that in count is there a way to reuse shadow maps? (like the new unity 5.1 "shared shadowmaps") What about "SHADOW_NODE_REUSE" does this helps with VR? how do I enable this feature in my shadow node?
3. Any other tips??
Thank you so much!!
-
frostbyte
- Orc Shaman
- Posts: 737
- Joined: Fri May 31, 2013 2:28 am
- x 65
Re: [2.1] Performance Optimizations for VR
port pczManager?way to avoid rendering those objects?
or maybe https://github.com/gigc/Janua
i imagine you're looking for some more advanced occlusion-culling method but anyway this old methods used to do the trick in such situations...
i rembmer dark_sylink mumbling something about adding occlution-culling in 2.1
but i think its a bit far ahead...
the woods are lovely dark and deep
but i have promises to keep
and miles to code before i sleep
and miles to code before i sleep..
coolest videos link( two minutes paper )...
https://www.youtube.com/user/keeroyz/videos
but i have promises to keep
and miles to code before i sleep
and miles to code before i sleep..
coolest videos link( two minutes paper )...
https://www.youtube.com/user/keeroyz/videos
-
al2950
- OGRE Expert User

- Posts: 1227
- Joined: Thu Dec 11, 2008 7:56 pm
- Location: Bristol, UK
- x 157
Re: [2.1] Performance Optimizations for VR
1. What you are talking about here is occlusion culling. When unity says 'single pass culling' I assume it is only talking about frustrum culling, and yes currently Ogre 2.x will cull the frustrum per pass. There are plans to allow sharing of culling data between passes, however the truth is Ogre is now so fast at frustrum culling it is no longer a hot spot in the rendering pipeline so the priority of that feature has been lowered.
However as I first stated what you need is occlusion culling. There is quite a lot of material on this on the net, but the most recent approach that I am aware of is to rasterise marked occluders on the CPU and then do culling checks against the rasterised scene and the bounding boxes of other objects. There are commercial libs available like http://umbra3d.com/, but there is also open source examples https://software.intel.com/en-us/blogs/ ... g-update-2
2. In short yes, however in the stereo example in the Ogre 2.1 samples it uses 2 separate workspaces, which as far as I am aware can not share nodes between them. If you used a single workspace you would do something like this
NB I tested the above code and found a bug to do with the viewport property, but it does work if the bug is fixed! I will create a pull request now
3. Probably! But it would be useful to know more about your rendering setup. There are a few things you can tweek if you understand how the new HLMS + AZDO works (not that I do!!)
Hope that helps
**EDIT** Pull request here; https://bitbucket.org/sinbad/ogre/pull- ... wport/diff
However as I first stated what you need is occlusion culling. There is quite a lot of material on this on the net, but the most recent approach that I am aware of is to rasterise marked occluders on the CPU and then do culling checks against the rasterised scene and the bounding boxes of other objects. There are commercial libs available like http://umbra3d.com/, but there is also open source examples https://software.intel.com/en-us/blogs/ ... g-update-2
2. In short yes, however in the stereo example in the Ogre 2.1 samples it uses 2 separate workspaces, which as far as I am aware can not share nodes between them. If you used a single workspace you would do something like this
Code: Select all
compositor_node StereoRenderingNode
{
in 0 rt_renderwindow
target rt_renderwindow
{
pass clear
{
colour_value 0.2 0.4 0.6 1
}
//left eye
pass render_scene
{
viewport 0.0 0.0 0.5 1.0
camera left
shadows shadowNode first
}
//right eye
pass render_scene
{
viewport 0.5 0.0 0.5 1.0
camera right
//Ogre will try and re-render shadow maps as its using a different camera
//So explicitly tell it to reuse
shadows shadowNode reuse
}
}
}
3. Probably! But it would be useful to know more about your rendering setup. There are a few things you can tweek if you understand how the new HLMS + AZDO works (not that I do!!)
Hope that helps
**EDIT** Pull request here; https://bitbucket.org/sinbad/ogre/pull- ... wport/diff
- dark_sylinc
- OGRE Team Member

- Posts: 5588
- Joined: Sat Jul 21, 2007 4:55 pm
- Location: Buenos Aires, Argentina
- x 1413
- Contact:
Re: [2.1] Performance Optimizations for VR
I think they mean this optimization I found about yesterday.al2950 wrote:When unity says 'single pass culling' I assume it is only talking about frustrum culling, and yes currently Ogre 2.x will cull the frustrum per pass.
Nothing major, quite clever; and if combined with GL_AMD_vertex_shader_viewport_index or VPAndRTArrayIndexFromAnyShaderFeedingRasterizerSupportedWithoutGSEmulation it can make for major CPU side optimizations (cull & rasterize to both eyes in one pass)
However clearly OP is GPU bound.
SW Occlusion culling (ala Intel's) is planned, but I have to admit they get put into the backburner over and over again, and considering OP's warehouse scenario it could greatly benefit him.
A Z Prepass could also help him. Out of the box Z Prepasses will soon be available in the compositor. But with a bit of hacking or cleverness it may be feasible for him today (render directly to a PF_D32_FLOAT view that shares the same depth buffer pool as the main render target, then render to that RTT; and set the RTT's depth buffer format to the same as the depth buffer view, an example is in the manual):
Code: Select all
compositor_node Example2_fixed
{
//Instruct we want to use a depth texture (32-bit float). The “depth_texture” keyword is necessary. Specifying The depth format is optional and so is the depth pool. However recommended to specify them to avoid surprises.
texture rt0 target_width target_height PF_R8G8B8 depth_format PF_D32_FLOAT depth_texture depth_pool 1
//Declare the depth texture view (which becomes so by using PF_D32_FLOAT as format). Settings MUST match (depth format, pools, resolution). Specifying the depth pool is necessary, otherwise the depth texture will get its own depth buffer, instead of becoming a view.
texture depthTexture target_width target_height PF_D32_FLOAT depth_pool 1
//Z prepass
target depthTexture
{
pass clear {}
pass render_scene
{
}
}
//Real pass
target rt0
{
pass clear {}
pass render_scene
{
shadows myShadowNode
}
}
}Forcing some shadow mapping reuse will certainly improve framerates, but you might get some artifacts in one of the eyes (the chance is extremely minor if the scene is physically big enough). I should consider a way to address this so that reuse is maximized for VR by taking into account the two cameras.
Make sure your depth shadow maps are of format PF_D32_FLOAT or PF_D16_FLOAT. Lowering the resolution can help.
Also, what GPU do you have and what framerate do you get (avg, max, min)? There is a reason Oculus is recommending monstrous machines for a good experience.
Done!**EDIT** Pull request here; https://bitbucket.org/sinbad/ogre/pull- ... wport/diff
-
xrgo
- OGRE Expert User

- Posts: 1148
- Joined: Sat Jul 06, 2013 10:59 pm
- Location: Chile
- x 169
Re: [2.1] Performance Optimizations for VR
Thank you so much everyone for all the information
I already tried al2950's workspace for reuse shadow maps and helped a lot!!! thank you so much!!
next I am going to try dark_sylinc's Z prepass =)
I am working on my laptop that has a 540m, with the reuse of shadowmaps I jump from like 9fps (yes, 9, my scene is quite heavy, and my laptop quite old) to 16fps!
And we have a machine here that has a 970, with that we jump from 70 to 110 fps!
All with the two viewports, distortion, some simple post process effects, and lots of objects, polys, physics, etc, and looking at the warehouse that is the direction with lower fps, in some directions now I getting like 200 fps. And I don't see any shadow glitch at all!
Thanks everyone!
I already tried al2950's workspace for reuse shadow maps and helped a lot!!! thank you so much!!
next I am going to try dark_sylinc's Z prepass =)
I am working on my laptop that has a 540m, with the reuse of shadowmaps I jump from like 9fps (yes, 9, my scene is quite heavy, and my laptop quite old) to 16fps!
And we have a machine here that has a 970, with that we jump from 70 to 110 fps!
All with the two viewports, distortion, some simple post process effects, and lots of objects, polys, physics, etc, and looking at the warehouse that is the direction with lower fps, in some directions now I getting like 200 fps. And I don't see any shadow glitch at all!
Thanks everyone!
- TaaTT4
- OGRE Contributor

- Posts: 267
- Joined: Wed Apr 23, 2014 3:49 pm
- Location: Bologna, Italy
- x 75
- Contact:
Re: [2.1] Performance Optimizations for VR
Hi guys,
I'm also working on adding VR support to my game.
Unlike what @xrgo is doing, I'll use OpenVR SDK.
Anyway, both SDK should work in a very similar way so the optimization strategies should be valid regardless of the SDK chosen.
Leaving apart the Z-buffer prepass technique which isn't strictly related to VR (and of which I have to investigate if it will bring benefits to my scene), what other optimization tricks can be used when you render for VR?
Sharing shadows between the different eyes it's almost mandatory since it's one of the huge task in the rendering pipeline.
Is it still true that the sharing can happen just in the same workspace?
What other can be done?
OT
What the hell were you thinking about, Microsoft?
Best thing is they have bothered to shorten the words viewport, render target and geometry shader.
I'm also working on adding VR support to my game.
Unlike what @xrgo is doing, I'll use OpenVR SDK.
Anyway, both SDK should work in a very similar way so the optimization strategies should be valid regardless of the SDK chosen.
Leaving apart the Z-buffer prepass technique which isn't strictly related to VR (and of which I have to investigate if it will bring benefits to my scene), what other optimization tricks can be used when you render for VR?
Sharing shadows between the different eyes it's almost mandatory since it's one of the huge task in the rendering pipeline.
Is it still true that the sharing can happen just in the same workspace?
What other can be done?
OT
By far, best variable name EVER!dark_sylinc wrote: VPAndRTArrayIndexFromAnyShaderFeedingRasterizerSupportedWithoutGSEmulation
What the hell were you thinking about, Microsoft?
Best thing is they have bothered to shorten the words viewport, render target and geometry shader.
Senior programmer at 505 Games; former senior engine programmer at Sandbox Games
Worked on: Racecraft Esport — Racecraft Coin-Op, Victory: The Age of Racing
- dark_sylinc
- OGRE Team Member

- Posts: 5588
- Joined: Sat Jul 21, 2007 4:55 pm
- Location: Buenos Aires, Argentina
- x 1413
- Contact:
Re: [2.1] Performance Optimizations for VR
TaaTT4 is in a slightly different boat than xrgo because xrgo is GPU bottlenecked, but TaaTT4 is in a rare situation where you're CPU bottlenecked.
There is a CPU side VR optimization we haven't implemented at Ogre, regarding frustum culling. It involves changing the camera (only for culling) and enlarging its frustum so that it encloses the two frustums from the two eyes:

The black frustums are the two eye's cameras. The blue frustum is the combined one that encloses both of them for frustum culling. The gains of this method are CPU side:
There is a CPU side VR optimization we haven't implemented at Ogre, regarding frustum culling. It involves changing the camera (only for culling) and enlarging its frustum so that it encloses the two frustums from the two eyes:

The black frustums are the two eye's cameras. The blue frustum is the combined one that encloses both of them for frustum culling. The gains of this method are CPU side:
- Frustum culling only needs to happen once instead of twice.
- The RenderQueue lists can be reused (i.e. no need to sort them again).
- Ideally we should be able to reuse the Command Buffer's commands while only changing the worldViewProj, view and viewProj matrices. But this could be difficult to achieve (quite hard, not impossible)
- Con: Some objects that should be culled off will still be rendered (wastes a little more GPU performance)
- TaaTT4
- OGRE Contributor

- Posts: 267
- Joined: Wed Apr 23, 2014 3:49 pm
- Location: Bologna, Italy
- x 75
- Contact:
Re: [2.1] Performance Optimizations for VR
The big frustum!dark_sylinc wrote: Ideally we should be able to reuse the Command Buffer's commands while only changing the worldViewProj, view and viewProj matrices. But this could be difficult to achieve (quite hard, not impossible)
I've seen this technique somewhere before in some development blogs, but implementing it seems a bit beyond my skills.
I've already changed OGRE source code a bit for my needs, but the modifications requested to achieve the enclosing frustum sound scary and very close to the metal.
Maybe I could give it a try in the future.
Senior programmer at 505 Games; former senior engine programmer at Sandbox Games
Worked on: Racecraft Esport — Racecraft Coin-Op, Victory: The Age of Racing