No
Yes
View More
View Less
Working...
Close
OK
Cancel
Confirm
System Message
Delete
Schedule
An unknown error has occurred and your request could not be completed. Please contact support.
Scheduled
Wait Listed
Personal Calendar
Speaking
Conference Event
Meeting
Interest
Schedule TBD
Conflict Found
This session is already scheduled at another time. Would you like to...
Loading...
Please enter a maximum of {0} characters.
Please enter a maximum of {0} words.
must be 50 characters or less.
must be 40 characters or less.
Session Summary
We were unable to load the map image.
This has not yet been assigned to a map.
Search Catalog
Reply
Replies ()
Search
New Post
Microblog
Microblog Thread
Post Reply
Post
Your session timed out.
This web page is not optimized for viewing on a mobile device. Visit this site in a desktop browser to access the full set of features.
2017 GTC San Jose

S7343 - Optimizer's Toolbox: Fast CUDA Techniques for Real-Time Image Processing

Session Speakers
Session Description

Take your kernels to the next level with performance-enhancing techniques for all levels of the CUDA memory hierarchy. We'll share lessons gleaned from implementing demanding image-processing algorithms into the real-time visual simulation world. From CPU prototype to optimized GPU implementation, one algorithm saw 150,000X speedup. Techniques to be presented include: instantaneous image decimation; CDF via warp shuffle; block and grid shapes for easy-to-program cache optimization; designing XY-separable kernels and their intermediate data; and sliding window tradeoffs for maximum cache locality. Straightforward examples will make these optimizations easy to add to your CUDA toolbox.


Additional Session Information
Intermediate
Talk
Performance Optimization Real-Time Graphics Video and Image Processing
Aerospace Defense Manufacturing Media & Entertainment Software
50 minutes
Session Schedule