mirror of
https://github.com/ceres-solver/ceres-solver.git
synced 2026-08-29 16:40:38 +08:00
bdee4d6172
Instead of pre-computing pemutation from block-sparse to CRS order, index of value in CRS matrix is computed in the process of updating values using block-sparse structure. When it is possible to update values via a simple host-to-device copy, block-sparse structure on GPU is discarded after computing CRS structure. Computing index is significantly slower than using pre-computed permutation, but is still hidden by host-to-device transfer. On problems from BAL dataset this results into reduction of extra gpu memory consumption from 33% (permutation stored as 32-bit indices) to ~10% for storing block-sparse structure. Benchmark results: ======================= CUDA Device Properties ====================== Cuda version : 11.8 Device ID : 0 Device name : NVIDIA GeForce RTX 2080 Ti Total GPU memory : 11012 MiB GPU memory available : 10852 MiB Compute capability : 7.5 Warp size : 32 Max threads per block: 1024 Max threads per dim : 1024 1024 64 Max grid size : 2147483647 65535 65535 Multiprocessor count : 68 ==================================================================== Running ./bin/evaluation_benchmark Run on (112 X 3200 MHz CPU s) CPU Caches: L1 Data 32 KiB (x56) L1 Instruction 32 KiB (x56) L2 Unified 1024 KiB (x56) L3 Unified 39424 KiB (x2) Load Average: 24.58, 11.75, 8.52 ----------------------------------------------------------------------- Benchmark Time ----------------------------------------------------------------------- Using on-the-fly computation of CRS index corresponding to block-sparse index: JacobianToCRS<g/final/problem-4585-1324582-pre.txt> 1607 ms JacobianToCRSView<g/final/problem-4585-1324582-pre.txt> 564 ms JacobianToCRSMatrix<g/final/problem-4585-1324582-pre.txt> 2226 ms JacobianToCRSViewUpdate<g/final/problem-4585-1324582-pre.txt> 228 ms JacobianToCRSMatrixUpdate<g/final/problem-4585-1324582-pre.txt> 400 ms Using precomputed permutation: JacobianToCRS</final/problem-4585-1324582-pre.txt> 1656 ms JacobianToCRSView</final/problem-4585-1324582-pre.txt> 553 ms JacobianToCRSMatrix</final/problem-4585-1324582-pre.txt> 2255 ms JacobianToCRSViewUpdate</final/problem-4585-1324582-pre.txt> 228 ms JacobianToCRSMatrixUpdate</final/problem-4585-1324582-pre.txt> 406 ms Performance of JacobianToCRSViewUpdate is still limited by host-to-device transfer, and JacobianToCRSView is faster than computing CRS structure on CPU. Change-Id: Ifb6910fb01ae6071400d36c277846fadc5857964
114 lines
3.6 KiB
YAML
114 lines
3.6 KiB
YAML
name: Linux
|
|
|
|
on: [push, pull_request]
|
|
|
|
jobs:
|
|
build:
|
|
name: ${{matrix.os}}-${{matrix.build_type}}-${{matrix.lib}}-${{matrix.gpu}}
|
|
runs-on: ubuntu-latest
|
|
container: ${{matrix.os}}
|
|
defaults:
|
|
run:
|
|
shell: bash -e -o pipefail {0}
|
|
env:
|
|
CCACHE_DIR: ${{github.workspace}}/ccache
|
|
CMAKE_GENERATOR: Ninja
|
|
DEBIAN_FRONTEND: noninteractive
|
|
strategy:
|
|
fail-fast: true
|
|
matrix:
|
|
os:
|
|
- ubuntu:20.04
|
|
- ubuntu:22.04
|
|
build_type:
|
|
- Release
|
|
lib:
|
|
- shared
|
|
- static
|
|
gpu:
|
|
- cuda
|
|
- no-cuda
|
|
|
|
steps:
|
|
- uses: actions/checkout@v3
|
|
|
|
- name: Setup Dependencies
|
|
run: |
|
|
apt-get update
|
|
apt-get install -y \
|
|
build-essential \
|
|
ccache \
|
|
cmake \
|
|
libbenchmark-dev \
|
|
libblas-dev \
|
|
libeigen3-dev \
|
|
libgflags-dev \
|
|
libgoogle-glog-dev \
|
|
liblapack-dev \
|
|
libmetis-dev \
|
|
libsuitesparse-dev \
|
|
ninja-build \
|
|
wget
|
|
# nvidia cuda toolkit shipped with 20.04 LTS does not support stream-ordered allocations
|
|
- name: Setup CUDA toolkit repositories (20.04)
|
|
if: matrix.gpu == 'cuda' && matrix.os == 'ubuntu:20.04'
|
|
run: |
|
|
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/cuda-keyring_1.0-1_all.deb
|
|
dpkg -i cuda-keyring_1.0-1_all.deb
|
|
# nvidia cuda toolkit + gcc combo shipped with 22.04LTS is broken
|
|
# and is not able to compile code that uses thrust
|
|
# https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1006962
|
|
- name: Setup CUDA toolkit repositories (22.04)
|
|
if: matrix.gpu == 'cuda' && matrix.os == 'ubuntu:22.04'
|
|
run: |
|
|
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2204/x86_64/cuda-keyring_1.0-1_all.deb
|
|
dpkg -i cuda-keyring_1.0-1_all.deb
|
|
- name: Setup CUDA toolkit
|
|
if: matrix.gpu == 'cuda'
|
|
run: |
|
|
apt-get update
|
|
apt-get install -y cuda
|
|
echo "CUDACXX=/usr/local/cuda/bin/nvcc" >> $GITHUB_ENV
|
|
|
|
- name: Cache Build
|
|
id: cache-build
|
|
uses: actions/cache@v3
|
|
with:
|
|
path: ${{env.CCACHE_DIR}}
|
|
key: ${{matrix.os}}-ccache-${{matrix.build_type}}-${{matrix.lib}}-${{matrix.gpu}}-${{github.run_id}}
|
|
restore-keys: ${{matrix.os}}-ccache-${{matrix.build_type}}-${{matrix.lib}}-${{matrix.gpu}}-
|
|
|
|
- name: Setup Environment
|
|
if: matrix.build_type == 'Release'
|
|
run: |
|
|
echo 'CXXFLAGS=-flto' >> $GITHUB_ENV
|
|
|
|
- name: Configure
|
|
run: |
|
|
cmake -S . -B build_${{matrix.build_type}} \
|
|
-DBUILD_SHARED_LIBS=${{matrix.lib == 'shared'}} \
|
|
-DUSE_CUDA=${{matrix.gpu == 'cuda'}} \
|
|
-DCMAKE_BUILD_TYPE=${{matrix.build_type}} \
|
|
-DCMAKE_C_COMPILER_LAUNCHER=$(which ccache) \
|
|
-DCMAKE_CXX_COMPILER_LAUNCHER=$(which ccache) \
|
|
-DCMAKE_INSTALL_PREFIX=${{github.workspace}}/install
|
|
|
|
- name: Build
|
|
run: |
|
|
cmake --build build_${{matrix.build_type}} \
|
|
--config ${{matrix.build_type}}
|
|
|
|
- name: Test
|
|
if: matrix.gpu == 'no-cuda'
|
|
run: |
|
|
cd build_${{matrix.build_type}}/
|
|
ctest --config ${{matrix.build_type}} \
|
|
--output-on-failure \
|
|
-j$(nproc)
|
|
|
|
- name: Install
|
|
run: |
|
|
cmake --build build_${{matrix.build_type}}/ \
|
|
--config ${{matrix.build_type}} \
|
|
--target install
|