My recent set of optimizations, has shown a hot spot around forward+ collectLights functions. This is compounded on my system as I have a number of scene passes in a single compositor, not mention multiple cameras and renderwindows. Currently each call to ForwardPlus::CollectLights takes around 0.5ms, which in my system adds up to over 3ms, and sometimes more. I have a 11ms budget so that is killing me at the moment.
Anyway, this looks like an obvious thing to put into a compute shader, so I was wondering if anyone has tried or if anyone knows of a reason why it could not be converted to a compute shader?
[2.1+] Can we put forward+ 'collectLights' funcs, into a Compute shader?
-
al2950
- OGRE Expert User

- Posts: 1227
- Joined: Thu Dec 11, 2008 7:56 pm
- Location: Bristol, UK
- x 157
- dark_sylinc
- OGRE Team Member

- Posts: 5588
- Joined: Sat Jul 21, 2007 4:55 pm
- Location: Buenos Aires, Argentina
- x 1413
- Contact:
Re: [2.1+] Can we put forward+ 'collectLights' funcs, into a Compute shader?
It is possible yes. Console games do that because they're often CPU bound (PS4 and XBox One have a very weak CPU).
Because Ogre is usually GPU bound on PC, collecting lights on the CPU is what makes most sense for us.
I don't know if that will solve your problem though, as you're shifting the problem to a different chip and 0.5ms is not actually much. Your problem is the algorithm running multiple times. If it's because of VR, I don't know if it could be optimized in a way data could be reused at the expense of some extra overhead.
There's data that could be reused or optimizations specific for your workload:
For example ForwardClustered::collectLights calls mSceneManager->cullLights which collects the lights that are intersecting the camera frustum, and then frustum culls in multiple threads against the subdivided minifrustums (the mSceneManager->cullLights it acts as a sort of broadphase filter). This implies two things:
Remember you can turn off F+ for specific passes via script:
Another issue is that theoretically all of those light collection could be done in the background. Right Ogre will perform:
Whereas for your specific workload the ideal scenario would be:
While in the future we want to support such schemes, it is not currently possible (unless you do it by hand?)
I am also evaluating a different approach at ForwardClustered (also CPU based) that hopefully will be faster, but that's not yet close to getting implemented.
Cheers
Because Ogre is usually GPU bound on PC, collecting lights on the CPU is what makes most sense for us.
I don't know if that will solve your problem though, as you're shifting the problem to a different chip and 0.5ms is not actually much. Your problem is the algorithm running multiple times. If it's because of VR, I don't know if it could be optimized in a way data could be reused at the expense of some extra overhead.
There's data that could be reused or optimizations specific for your workload:
For example ForwardClustered::collectLights calls mSceneManager->cullLights which collects the lights that are intersecting the camera frustum, and then frustum culls in multiple threads against the subdivided minifrustums (the mSceneManager->cullLights it acts as a sort of broadphase filter). This implies two things:
- That mSceneManager->cullLights saves significant time later on the individual chunks. This isn't always true. For example if 95% of the lights are always intersecting the camera frustum, we're wasting our cycles.
- This could be changed in a way that subsequent passes reuse those results, assuming the previous pass includes all the potential light candidates the current pass will need. No need to call mSceneManager->cullLights again.
Remember you can turn off F+ for specific passes via script:
Code: Select all
enable_forwardplus falseCode: Select all
1st pass collectLights in multiple threads
sync
1st pass render
2nd pass collectLights in multiple threads
sync
2nd pass render
Nth pass collectLights in multiple threads
sync
Nth pass renderCode: Select all
1st pass collectLights
2nd pass collectLights
Nth pass collectLights
1st pass sync
1st pass render
2nd pass sync
2nd pass render
3rd pass sync
3rd pass renderI am also evaluating a different approach at ForwardClustered (also CPU based) that hopefully will be faster, but that's not yet close to getting implemented.
Cheers
-
al2950
- OGRE Expert User

- Posts: 1227
- Joined: Thu Dec 11, 2008 7:56 pm
- Location: Bristol, UK
- x 157
Re: [2.1+] Can we put forward+ 'collectLights' funcs, into a Compute shader?
Thank you very much for the detailed reply
I agree, to a point. However if you are updating reflection probes which actually only draw a tiny bit of the scene, and so GPU time is actually very small, then 0.5ms is a huge amount (in my case 50% of a single scene pass is spent in collectLights!). My view for this use case is that the GPU should dramatically reduce the total time for a scene pass is collect lights was executed on the GPU as a compute shader. From an Ogre point of view this could be an optional thingy, or even perhaps a different ForwardPlus algorithm.dark_sylinc wrote: Mon Feb 05, 2018 5:20 pm as you're shifting the problem to a different chip and 0.5ms is not actually much
Interesting point, I already do this for shadows and scene culling, I was not sure if the Forward+ implementations (maths wise) was tide very specifically to the camera frustrum. If not I might be able to put the collectLight call into _cullPhase01 and use the culling camera. Ill give this a go tomorrowdark_sylinc wrote: Mon Feb 05, 2018 5:20 pm For VR, the obvious optimization is that mSceneManager->cullLights should use the "frustum that encloses the two frustum" optimization and then reuse those results for the 2nd pass.
Yes, I do already use this, but currently have transparent objects rendered in a separate pass which have F+ enabled, this is probably a bit of waste! So I shold be able to save some time there.dark_sylinc wrote: Mon Feb 05, 2018 5:20 pm Remember you can turn off F+ for specific passes via script: