[2.1+] Can we put forward+ 'collectLights' funcs, into a Compute shader?

Discussion area about developing with Ogre-Next (2.1, 2.2 and beyond)


Post Reply
al2950
OGRE Expert User
OGRE Expert User
Posts: 1227
Joined: Thu Dec 11, 2008 7:56 pm
Location: Bristol, UK
x 157

[2.1+] Can we put forward+ 'collectLights' funcs, into a Compute shader?

Post by al2950 »

My recent set of optimizations, has shown a hot spot around forward+ collectLights functions. This is compounded on my system as I have a number of scene passes in a single compositor, not mention multiple cameras and renderwindows. Currently each call to ForwardPlus::CollectLights takes around 0.5ms, which in my system adds up to over 3ms, and sometimes more. I have a 11ms budget so that is killing me at the moment.

Anyway, this looks like an obvious thing to put into a compute shader, so I was wondering if anyone has tried or if anyone knows of a reason why it could not be converted to a compute shader?
User avatar
dark_sylinc
OGRE Team Member
OGRE Team Member
Posts: 5588
Joined: Sat Jul 21, 2007 4:55 pm
Location: Buenos Aires, Argentina
x 1413
Contact:

Re: [2.1+] Can we put forward+ 'collectLights' funcs, into a Compute shader?

Post by dark_sylinc »

It is possible yes. Console games do that because they're often CPU bound (PS4 and XBox One have a very weak CPU).
Because Ogre is usually GPU bound on PC, collecting lights on the CPU is what makes most sense for us.

I don't know if that will solve your problem though, as you're shifting the problem to a different chip and 0.5ms is not actually much. Your problem is the algorithm running multiple times. If it's because of VR, I don't know if it could be optimized in a way data could be reused at the expense of some extra overhead.

There's data that could be reused or optimizations specific for your workload:

For example ForwardClustered::collectLights calls mSceneManager->cullLights which collects the lights that are intersecting the camera frustum, and then frustum culls in multiple threads against the subdivided minifrustums (the mSceneManager->cullLights it acts as a sort of broadphase filter). This implies two things:
  1. That mSceneManager->cullLights saves significant time later on the individual chunks. This isn't always true. For example if 95% of the lights are always intersecting the camera frustum, we're wasting our cycles.
  2. This could be changed in a way that subsequent passes reuse those results, assuming the previous pass includes all the potential light candidates the current pass will need. No need to call mSceneManager->cullLights again.
For VR, the obvious optimization is that mSceneManager->cullLights should use the "frustum that encloses the two frustum" optimization and then reuse those results for the 2nd pass.

Remember you can turn off F+ for specific passes via script:

Code: Select all

enable_forwardplus false
Another issue is that theoretically all of those light collection could be done in the background. Right Ogre will perform:

Code: Select all

1st pass collectLights in multiple threads
sync
1st pass render
2nd pass collectLights in multiple threads
sync
2nd pass render
Nth pass collectLights in multiple threads
sync
Nth pass render
Whereas for your specific workload the ideal scenario would be:

Code: Select all

1st pass collectLights
2nd pass collectLights
Nth pass collectLights
1st pass sync
1st pass render
2nd pass sync
2nd pass render
3rd pass sync
3rd pass render
While in the future we want to support such schemes, it is not currently possible (unless you do it by hand?)

I am also evaluating a different approach at ForwardClustered (also CPU based) that hopefully will be faster, but that's not yet close to getting implemented.

Cheers
al2950
OGRE Expert User
OGRE Expert User
Posts: 1227
Joined: Thu Dec 11, 2008 7:56 pm
Location: Bristol, UK
x 157

Re: [2.1+] Can we put forward+ 'collectLights' funcs, into a Compute shader?

Post by al2950 »

Thank you very much for the detailed reply
dark_sylinc wrote: Mon Feb 05, 2018 5:20 pm as you're shifting the problem to a different chip and 0.5ms is not actually much
I agree, to a point. However if you are updating reflection probes which actually only draw a tiny bit of the scene, and so GPU time is actually very small, then 0.5ms is a huge amount (in my case 50% of a single scene pass is spent in collectLights!). My view for this use case is that the GPU should dramatically reduce the total time for a scene pass is collect lights was executed on the GPU as a compute shader. From an Ogre point of view this could be an optional thingy, or even perhaps a different ForwardPlus algorithm.
dark_sylinc wrote: Mon Feb 05, 2018 5:20 pm For VR, the obvious optimization is that mSceneManager->cullLights should use the "frustum that encloses the two frustum" optimization and then reuse those results for the 2nd pass.
Interesting point, I already do this for shadows and scene culling, I was not sure if the Forward+ implementations (maths wise) was tide very specifically to the camera frustrum. If not I might be able to put the collectLight call into _cullPhase01 and use the culling camera. Ill give this a go tomorrow
dark_sylinc wrote: Mon Feb 05, 2018 5:20 pm Remember you can turn off F+ for specific passes via script:
Yes, I do already use this, but currently have transparent objects rendered in a separate pass which have F+ enabled, this is probably a bit of waste! So I shold be able to save some time there.
Post Reply