cuda 大师班

3.2万
27
2020-02-22 17:52:51
324
137
2081
98
https://www.udemy.com/course/cuda-programming-masterclass/ 课件和代码在https://github.com/zhangasia/cuda_masterclass.git
茶飘香,酒罢去,聚朋友,再回楼
视频选集
(1/83)
001 Very very important
07:50
002 Introduction to parallel programming
08:51
003 Parallel computing and Super computing
07:20
004 How to install CUDA toolkit and first look at CUDA program
06:13
005 Basic elements of CUDA program
16:51
006 Organization of threads in a CUDA program - threadIdx
08:39
007 Organization of thread in a CUDA program - blockIdx,blockDim,gridDim
06:15
008 Programming exercise 1
00:30
009 Unique index calculation using threadIdx blockId and blockDim
09:21
010 Unique index calculation for 2D grid 1
05:54
011 Unique index calculation for 2D grid 2
05:11
012 Memory transfer between host and device
11:14
013 Programming exercise 2
01:05
014 Sum array example with validity check
09:14
015 Sum array example with error handling
04:34
016 Sum array example with timing
08:19
017 Device properties
05:31
018 Summary
04:19
019 Understand the device better
08:47
020 All about warps
09:44
021 Warp divergence
12:29
022 Resource partitioning and latency hiding 1
05:36
023 Resource partitioning and latency hiding 2
10:42
024 Occupancy
11:17
025 Profile driven optimization with nvprof
12:05
026 Parallel reduction as synchronization example
19:09
027 Parallel reduction as warp divergence example
10:12
028 Parallel reduction with loop unrolling
07:04
029 Parallel reduction as warp unrolling
06:49
030 Reduction with complete unrolling
04:10
031 Performance comparison of reduction kernels
05:19
032 CUDA Dynamic parallelism
10:04
033 Reduction with dynamic parallelism
05:34
034 Summary
04:37
035 CUDA memory model
06:50
036 Different memory types in CUDA
09:05
037 Memory management and pinned memory
07:20
038 Zero copy memory
08:46
039 Unified memory
04:40
040 Global memory access patterns
12:56
041 Global memory writes
03:54
042 AOS vs SOA
06:04
043 Matrix transpose
19:35
044 Matrix transpose with unrolling
06:22
045 Matrix transpose with diagonal coordinate system
08:37
046 Summary
03:01
047 Introduction to CUDA shared memory
09:05
048 Shared memory access modes and memory banks
09:07
049 Row major and Column major access to shared memory
08:52
050 Static and Dynamic shared memory
04:20
051 Shared memory padding
05:45
052 Parallel reduction with shared memory
04:45
053 Synchronization in CUDA
03:39
054 Matrix transpose with shared memory
11:55
055 CUDA constant memory
13:11
056 Matrix transpose with Shared memory padding
05:49
057 CUDA warp shuffle instructions
15:00
058 Parallel reduction with warp shuffle instructions
03:51
059 Summary
02:11
060 Introduction to CUDA streams and events
06:26
061 How to use CUDA asynchronous functions
07:11
062 How to use CUDA streams
10:29
063 Overlapping memory transfer and kernel execution
05:24
064 Stream synchronization and blocking behavious of NULL stream
06:58
065 Explicit and implicit synchronization
02:32
066 CUDA events and timing with CUDA events
06:04
067 Creating inter stream dependencies with events
04:32
068 Introduction to different types of instructions in CUDA
04:02
069 Floating point operations
06:47
070 Standard and Instrict functions
08:30
071 Atomic functions
08:24
072 Scan algorithm introduction
05:39
073 Simple parallel scan
08:25
074 Work efficient parallel exclusive scan
09:34
075 Work efficient parallel inclusive scan
07:42
076 Parallel scan for large data sets
04:53
077 Parallel Compact algorithm
07:50
078 Introduction part 1
08:06
079 Introduction part 2
11:42
080 Digital image processing
09:40
081 Digital image fundametals _ Human perception
11:11
082 Digital image fundamentals _ Image formation
15:23
083 OpenCV installation
06:29
客服
顶部
赛事库 课堂 2021拜年纪