mirror of
https://github.com/ceres-solver/ceres-solver.git
synced 2026-08-29 08:34:37 +08:00
Block-sparse to CRS conversion using block-structure
Instead of pre-computing pemutation from block-sparse to CRS order, index of value in CRS matrix is computed in the process of updating values using block-sparse structure. When it is possible to update values via a simple host-to-device copy, block-sparse structure on GPU is discarded after computing CRS structure. Computing index is significantly slower than using pre-computed permutation, but is still hidden by host-to-device transfer. On problems from BAL dataset this results into reduction of extra gpu memory consumption from 33% (permutation stored as 32-bit indices) to ~10% for storing block-sparse structure. Benchmark results: ======================= CUDA Device Properties ====================== Cuda version : 11.8 Device ID : 0 Device name : NVIDIA GeForce RTX 2080 Ti Total GPU memory : 11012 MiB GPU memory available : 10852 MiB Compute capability : 7.5 Warp size : 32 Max threads per block: 1024 Max threads per dim : 1024 1024 64 Max grid size : 2147483647 65535 65535 Multiprocessor count : 68 ==================================================================== Running ./bin/evaluation_benchmark Run on (112 X 3200 MHz CPU s) CPU Caches: L1 Data 32 KiB (x56) L1 Instruction 32 KiB (x56) L2 Unified 1024 KiB (x56) L3 Unified 39424 KiB (x2) Load Average: 24.58, 11.75, 8.52 ----------------------------------------------------------------------- Benchmark Time ----------------------------------------------------------------------- Using on-the-fly computation of CRS index corresponding to block-sparse index: JacobianToCRS<g/final/problem-4585-1324582-pre.txt> 1607 ms JacobianToCRSView<g/final/problem-4585-1324582-pre.txt> 564 ms JacobianToCRSMatrix<g/final/problem-4585-1324582-pre.txt> 2226 ms JacobianToCRSViewUpdate<g/final/problem-4585-1324582-pre.txt> 228 ms JacobianToCRSMatrixUpdate<g/final/problem-4585-1324582-pre.txt> 400 ms Using precomputed permutation: JacobianToCRS</final/problem-4585-1324582-pre.txt> 1656 ms JacobianToCRSView</final/problem-4585-1324582-pre.txt> 553 ms JacobianToCRSMatrix</final/problem-4585-1324582-pre.txt> 2255 ms JacobianToCRSViewUpdate</final/problem-4585-1324582-pre.txt> 228 ms JacobianToCRSMatrixUpdate</final/problem-4585-1324582-pre.txt> 406 ms Performance of JacobianToCRSViewUpdate is still limited by host-to-device transfer, and JacobianToCRSView is faster than computing CRS structure on CPU. Change-Id: Ifb6910fb01ae6071400d36c277846fadc5857964
This commit is contained in:
@@ -49,20 +49,23 @@ jobs:
|
||||
libsuitesparse-dev \
|
||||
ninja-build \
|
||||
wget
|
||||
- name: Setup CUDA toolkit (system repositories)
|
||||
# nvidia cuda toolkit shipped with 20.04 LTS does not support stream-ordered allocations
|
||||
- name: Setup CUDA toolkit repositories (20.04)
|
||||
if: matrix.gpu == 'cuda' && matrix.os == 'ubuntu:20.04'
|
||||
run: |
|
||||
apt-get install -y \
|
||||
nvidia-cuda-dev \
|
||||
nvidia-cuda-toolkit
|
||||
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/cuda-keyring_1.0-1_all.deb
|
||||
dpkg -i cuda-keyring_1.0-1_all.deb
|
||||
# nvidia cuda toolkit + gcc combo shipped with 22.04LTS is broken
|
||||
# and is not able to compile code that uses thrust
|
||||
# https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1006962
|
||||
- name: Setup CUDA toolkit (nvidia repositories)
|
||||
- name: Setup CUDA toolkit repositories (22.04)
|
||||
if: matrix.gpu == 'cuda' && matrix.os == 'ubuntu:22.04'
|
||||
run: |
|
||||
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.0-1_all.deb
|
||||
dpkg -i cuda-keyring_1.0-1_all.deb
|
||||
- name: Setup CUDA toolkit
|
||||
if: matrix.gpu == 'cuda'
|
||||
run: |
|
||||
apt-get update
|
||||
apt-get install -y cuda
|
||||
echo "CUDACXX=/usr/local/cuda/bin/nvcc" >> $GITHUB_ENV
|
||||
|
||||
Reference in New Issue
Block a user