Remove "using ceres:foo" directives from example code. The using
directives actually make the code harder to read unless you already
know the ceres API. By making the namespace explicit it is clear
to the reader that these are functions and objects from the Ceres
API.
Change-Id: I89b1281c754bf71c0f82e39e1607c5e40a148388
cuda-memcheck has been deprecated and these tests will be reinstated
once we move to compute sanitizer.
Change-Id: I7e01cffdd00ffb7404dfef2a171ebaeca017cda8
1. Add a version history
2. Update copyright years across the code base
3. Run format_all.sh
4. Update version strings from 2.1.0 to 2.2.0 in the docs and
elsewhere.
Change-Id: I46d8d479d54bd6002d532785e67342106e73c9ac
* Make sphinx_rtd_theme a find module component to avoid hard-wiring it
into the module and allowing to report the theme in case it is missing
using the standard CMake package mechanism.
* Adjust find module cache variables names case to match the find module
name.
* Also report sphinx-build version for completeness.
* Invoke the find module only once. Calling find_package on the same
module is not needed.
Change-Id: I9d1bf0fcc0d44b9b37e624128812f348c5442ada
Given we no longer support Ubuntu 18.04 due to packaged GCC lacking
C++17 support we can bump the minimum required CMake version to the one
provided by Ubuntu 20.04 which is CMake 3.16. Consequently, this allows
to drop some of the legacy CMake logic.
Change-Id: I1f05d4c5681d10aa7faa0800ef4a803be2f5b7dd
AppleClang 14.0.0.14000029 warns about a potential security problem
while invoking the sprintf C function:
internal/ceres/fixed_array_test.cc:469:3: warning: 'sprintf' is deprecated: This function is provided for compatibility reasons only. Due to security concerns inherent in the design of sprintf(3), it is highly recommended that you use snprintf(3) instead. [-Wdeprecated-declarations]
sprintf(buf.data(), "foo"); // NOLINT(runtime/printf)
^
/Applications/Xcode_14.2.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX13.1.sdk/usr/include/stdio.h:188:1: note: 'sprintf' has been explicitly marked deprecated here
__deprecated_msg("This function is provided for compatibility reasons only. Due to security concerns inherent in the design of sprintf(3), it is highly recommended that you use snprintf(3) instead.")
^
/Applications/Xcode_14.2.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX13.1.sdk/usr/include/sys/cdefs.h:215:48: note: expanded from macro '__deprecated_msg'
#define __deprecated_msg(_msg) __attribute__((__deprecated__(_msg)))
Replace sprintf by snprintf to avoid this deprecation warning.
Change-Id: I6870c0bd4e390388d1d7bcec082cee272b234eba
This avoids the following CMake error when exporting Ceres build
directory to local CMake package registry:
export called with target "ceres" which requires target
"ceres_cuda_kernels" that is not in any export set.
Fixes#966
Change-Id: I5a75191fc414a3f138b19cb9d5850f7330f2d24a
Without this we start getting errors related to FindSphinx.cmake
I am not sure yet, what version of cmake we can assume in the wild
but the current minimum version supports the old behaviour and this
policy seems relatively recent (CMake version 3.27) so we should
have this workaround till we update our minimum required version.
Fixes https://github.com/ceres-solver/ceres-solver/issues/1002
Change-Id: I1beaac9ee27bc9ff85b64f53f72606a0424f2391
Converting fixed size vectors to dynamic ones allows to avoid
segmentation faults in Eigen's packet math if the corresponding
expressions are invoked within GMock matchers.
Fixes#996
Change-Id: I7da5599883825ab0e580678d3d55de19095b41b1
- Perform temporary buffer size estimation only once
- Allow construction from existing buffers with col/row structure
Change-Id: I73c291328f1e8ed9184aba5d7058df71cbc6a15d
1. In cuda_sparse_matrix.cc fix the order of fields in the initializer list.
2. Move a line of code to the ifdef branch which will use it.
Change-Id: If32ea14a287f845c1740e6f726c1007e86a4eeca
- In the top-level CMakeLists.txt, certain flags are passed to
disable warnings and increase the maximum size of an object file.
nvcc cannot handle these flags so we tell CMake to only use them
for C++ code (and not CUDA code).
- In internal/ceres/CMakeLists.txt, CMake is originally told to link
the import library cudart.lib when linking CUDA code. By default,
it seems that Visual Studio will link the static library
cudart_static.lib when linking CUDA code. So we avoid linking with
cudart.lib to avoid linking the same library twice.
Change-Id: I1fbf0d7e76d57b4338708757b27f5074722608cb
Detect when the number of non-zeros overflows when constructing
BlockSparseMatrix and CompressedRowSparseMatrix and return
with an error message instead of crashing.
Change-Id: I45e102f7c0519eef441ce0586b7adf96e4a954a9
Previously it could be the case that a residual block could return
a residual whose squared norm overflows and generates an infinity
which we did not detect. This would then lead to the trust region
minimizer incorrectly terminating indicating convergence while
generating a cost delta of NaN.
This change adds a check for that and also does two minor cosmetic
changes.
1. Reduce the level of nesting in program_evaluator.h by adding
an early return.
2. The error message when IterationZero fails now says that the
Initial residual and Jacobian failed, to indicate that the
optimizer had no chance to do any work.
Fixes https://github.com/ceres-solver/ceres-solver/issues/988
Thanks to @Ashray-g for reporting this.
Change-Id: I52ae7627a66f637135209dbb2e42935b52c8bc77
Converts BlockSparseMatrix into two instances of CudaSparseMatrix,
corresponding to left and right sub-matrix.
Values of submatrix E are always just copied as-is, and values of
submatrix F are copied if each row-block of F submatrix satisfies
at least one of the following conditions:
- There is atmost one cell in row-block
- Row block has height of 1 row
Otherwise, indices of values in CRS order corresponding to value indices
in block-sparse order are computed on-the-fly.
Change-Id: I14eee00c36ee74b6b83fc85927907641383abfc7
- Provides fall-back for older versions of CUDA toolkit
- Using older versions of CUDA toolkit might result in
over-synchronization
Change-Id: I545e6625d2342be30cb759b90bda379e555d7370
Instead of pre-computing pemutation from block-sparse to CRS order,
index of value in CRS matrix is computed in the process of updating
values using block-sparse structure.
When it is possible to update values via a simple host-to-device copy,
block-sparse structure on GPU is discarded after computing CRS
structure.
Computing index is significantly slower than using pre-computed
permutation, but is still hidden by host-to-device transfer.
On problems from BAL dataset this results into reduction of extra
gpu memory consumption from 33% (permutation stored as 32-bit indices)
to ~10% for storing block-sparse structure.
Benchmark results:
======================= CUDA Device Properties ======================
Cuda version : 11.8
Device ID : 0
Device name : NVIDIA GeForce RTX 2080 Ti
Total GPU memory : 11012 MiB
GPU memory available : 10852 MiB
Compute capability : 7.5
Warp size : 32
Max threads per block: 1024
Max threads per dim : 1024 1024 64
Max grid size : 2147483647 65535 65535
Multiprocessor count : 68
====================================================================
Running ./bin/evaluation_benchmark
Run on (112 X 3200 MHz CPU s)
CPU Caches:
L1 Data 32 KiB (x56)
L1 Instruction 32 KiB (x56)
L2 Unified 1024 KiB (x56)
L3 Unified 39424 KiB (x2)
Load Average: 24.58, 11.75, 8.52
-----------------------------------------------------------------------
Benchmark Time
-----------------------------------------------------------------------
Using on-the-fly computation of CRS index corresponding to block-sparse
index:
JacobianToCRS<g/final/problem-4585-1324582-pre.txt> 1607 ms
JacobianToCRSView<g/final/problem-4585-1324582-pre.txt> 564 ms
JacobianToCRSMatrix<g/final/problem-4585-1324582-pre.txt> 2226 ms
JacobianToCRSViewUpdate<g/final/problem-4585-1324582-pre.txt> 228 ms
JacobianToCRSMatrixUpdate<g/final/problem-4585-1324582-pre.txt> 400 ms
Using precomputed permutation:
JacobianToCRS</final/problem-4585-1324582-pre.txt> 1656 ms
JacobianToCRSView</final/problem-4585-1324582-pre.txt> 553 ms
JacobianToCRSMatrix</final/problem-4585-1324582-pre.txt> 2255 ms
JacobianToCRSViewUpdate</final/problem-4585-1324582-pre.txt> 228 ms
JacobianToCRSMatrixUpdate</final/problem-4585-1324582-pre.txt> 406 ms
Performance of JacobianToCRSViewUpdate is still limited by
host-to-device transfer, and JacobianToCRSView is faster than computing
CRS structure on CPU.
Change-Id: Ifb6910fb01ae6071400d36c277846fadc5857964
If using CUDA_SPARSE for an iterative solve on the GPU,
allocate the values array in BlockSparseMatrix to make copying
to the GPU faster.
Change-Id: I63c1d2512babd74fc275b277ac8c3eabf3ec1144
- TripletSparseMatrix in BlockRandomAccessSparseMatrix is replaced with
BlockSparseMatrix
- BlockSparseMatrix::ToCompressedRowSparseMatrix is performed in a
direct sort-less way
Change-Id: Ib951fda1b9394050e2c47a9721172c5e3c674801
1. Rename it to kRowShift.
2. Make it a static constexpr.
3. Change it to 2^32, which should allow for easier bit arithmetic
for the compiler than the previously used value of 10000000.
4. Change the name of the two associated private methods from
IntPairToLong to IntPairToInt64 and LongToIntPair to Int64ToIntPair.
Change-Id: I54d61bcf1121079b222ef518324de5cffc1be064