1. Add a version history
2. Update copyright years across the code base
3. Run format_all.sh
4. Update version strings from 2.1.0 to 2.2.0 in the docs and
elsewhere.
Change-Id: I46d8d479d54bd6002d532785e67342106e73c9ac
Given we no longer support Ubuntu 18.04 due to packaged GCC lacking
C++17 support we can bump the minimum required CMake version to the one
provided by Ubuntu 20.04 which is CMake 3.16. Consequently, this allows
to drop some of the legacy CMake logic.
Change-Id: I1f05d4c5681d10aa7faa0800ef4a803be2f5b7dd
AppleClang 14.0.0.14000029 warns about a potential security problem
while invoking the sprintf C function:
internal/ceres/fixed_array_test.cc:469:3: warning: 'sprintf' is deprecated: This function is provided for compatibility reasons only. Due to security concerns inherent in the design of sprintf(3), it is highly recommended that you use snprintf(3) instead. [-Wdeprecated-declarations]
sprintf(buf.data(), "foo"); // NOLINT(runtime/printf)
^
/Applications/Xcode_14.2.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX13.1.sdk/usr/include/stdio.h:188:1: note: 'sprintf' has been explicitly marked deprecated here
__deprecated_msg("This function is provided for compatibility reasons only. Due to security concerns inherent in the design of sprintf(3), it is highly recommended that you use snprintf(3) instead.")
^
/Applications/Xcode_14.2.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX13.1.sdk/usr/include/sys/cdefs.h:215:48: note: expanded from macro '__deprecated_msg'
#define __deprecated_msg(_msg) __attribute__((__deprecated__(_msg)))
Replace sprintf by snprintf to avoid this deprecation warning.
Change-Id: I6870c0bd4e390388d1d7bcec082cee272b234eba
- Perform temporary buffer size estimation only once
- Allow construction from existing buffers with col/row structure
Change-Id: I73c291328f1e8ed9184aba5d7058df71cbc6a15d
1. In cuda_sparse_matrix.cc fix the order of fields in the initializer list.
2. Move a line of code to the ifdef branch which will use it.
Change-Id: If32ea14a287f845c1740e6f726c1007e86a4eeca
Detect when the number of non-zeros overflows when constructing
BlockSparseMatrix and CompressedRowSparseMatrix and return
with an error message instead of crashing.
Change-Id: I45e102f7c0519eef441ce0586b7adf96e4a954a9
Previously it could be the case that a residual block could return
a residual whose squared norm overflows and generates an infinity
which we did not detect. This would then lead to the trust region
minimizer incorrectly terminating indicating convergence while
generating a cost delta of NaN.
This change adds a check for that and also does two minor cosmetic
changes.
1. Reduce the level of nesting in program_evaluator.h by adding
an early return.
2. The error message when IterationZero fails now says that the
Initial residual and Jacobian failed, to indicate that the
optimizer had no chance to do any work.
Fixes https://github.com/ceres-solver/ceres-solver/issues/988
Thanks to @Ashray-g for reporting this.
Change-Id: I52ae7627a66f637135209dbb2e42935b52c8bc77
Converts BlockSparseMatrix into two instances of CudaSparseMatrix,
corresponding to left and right sub-matrix.
Values of submatrix E are always just copied as-is, and values of
submatrix F are copied if each row-block of F submatrix satisfies
at least one of the following conditions:
- There is atmost one cell in row-block
- Row block has height of 1 row
Otherwise, indices of values in CRS order corresponding to value indices
in block-sparse order are computed on-the-fly.
Change-Id: I14eee00c36ee74b6b83fc85927907641383abfc7
- Provides fall-back for older versions of CUDA toolkit
- Using older versions of CUDA toolkit might result in
over-synchronization
Change-Id: I545e6625d2342be30cb759b90bda379e555d7370
Instead of pre-computing pemutation from block-sparse to CRS order,
index of value in CRS matrix is computed in the process of updating
values using block-sparse structure.
When it is possible to update values via a simple host-to-device copy,
block-sparse structure on GPU is discarded after computing CRS
structure.
Computing index is significantly slower than using pre-computed
permutation, but is still hidden by host-to-device transfer.
On problems from BAL dataset this results into reduction of extra
gpu memory consumption from 33% (permutation stored as 32-bit indices)
to ~10% for storing block-sparse structure.
Benchmark results:
======================= CUDA Device Properties ======================
Cuda version : 11.8
Device ID : 0
Device name : NVIDIA GeForce RTX 2080 Ti
Total GPU memory : 11012 MiB
GPU memory available : 10852 MiB
Compute capability : 7.5
Warp size : 32
Max threads per block: 1024
Max threads per dim : 1024 1024 64
Max grid size : 2147483647 65535 65535
Multiprocessor count : 68
====================================================================
Running ./bin/evaluation_benchmark
Run on (112 X 3200 MHz CPU s)
CPU Caches:
L1 Data 32 KiB (x56)
L1 Instruction 32 KiB (x56)
L2 Unified 1024 KiB (x56)
L3 Unified 39424 KiB (x2)
Load Average: 24.58, 11.75, 8.52
-----------------------------------------------------------------------
Benchmark Time
-----------------------------------------------------------------------
Using on-the-fly computation of CRS index corresponding to block-sparse
index:
JacobianToCRS<g/final/problem-4585-1324582-pre.txt> 1607 ms
JacobianToCRSView<g/final/problem-4585-1324582-pre.txt> 564 ms
JacobianToCRSMatrix<g/final/problem-4585-1324582-pre.txt> 2226 ms
JacobianToCRSViewUpdate<g/final/problem-4585-1324582-pre.txt> 228 ms
JacobianToCRSMatrixUpdate<g/final/problem-4585-1324582-pre.txt> 400 ms
Using precomputed permutation:
JacobianToCRS</final/problem-4585-1324582-pre.txt> 1656 ms
JacobianToCRSView</final/problem-4585-1324582-pre.txt> 553 ms
JacobianToCRSMatrix</final/problem-4585-1324582-pre.txt> 2255 ms
JacobianToCRSViewUpdate</final/problem-4585-1324582-pre.txt> 228 ms
JacobianToCRSMatrixUpdate</final/problem-4585-1324582-pre.txt> 406 ms
Performance of JacobianToCRSViewUpdate is still limited by
host-to-device transfer, and JacobianToCRSView is faster than computing
CRS structure on CPU.
Change-Id: Ifb6910fb01ae6071400d36c277846fadc5857964
If using CUDA_SPARSE for an iterative solve on the GPU,
allocate the values array in BlockSparseMatrix to make copying
to the GPU faster.
Change-Id: I63c1d2512babd74fc275b277ac8c3eabf3ec1144
- TripletSparseMatrix in BlockRandomAccessSparseMatrix is replaced with
BlockSparseMatrix
- BlockSparseMatrix::ToCompressedRowSparseMatrix is performed in a
direct sort-less way
Change-Id: Ib951fda1b9394050e2c47a9721172c5e3c674801
1. Rename it to kRowShift.
2. Make it a static constexpr.
3. Change it to 2^32, which should allow for easier bit arithmetic
for the compiler than the previously used value of 10000000.
4. Change the name of the two associated private methods from
IntPairToLong to IntPairToInt64 and LongToIntPair to Int64ToIntPair.
Change-Id: I54d61bcf1121079b222ef518324de5cffc1be064
In https://ceres-solver-review.git.corp.google.com/c/ceres-solver/+/23802
the computation of the norm of a quaternion
scale = 1/sqrt(q[0] * q[0] + q[1] * q[1] + q[2] * q[2] + q[3] * q[3]);
was replaced by
scale = 1/hypot(q[0], q[1], hypot(q[2], q[3]));
while this appear to be a more accurate computation because of the
use of hypot which can handle over and underflow it introduces a
bug for the case where q[2] = q[3] = 0.
While the hypot(q[2], q[3]) == 0 as scalars, if q[2] and q[3] are
jets, then the derivative will be NaN. Which means that even though
q[0] or q[1] is non-zero and the norm of the quaternion is non-zero,
and the resulting derivative is finite, this way of computing the
scale will produce nans in the derivative of scale.
The following quaternion will replicate the problem described above.
using Jet = ceres::Jet<double, 4>;
std::array<Jet, 4> quaternion = {Jet(1.0, 0), Jet(0.0, 1), Jet(0.0, 2), Jet(0.0, 3)};
This CL reverts the change to QuaternionRotatePoint and
adds a test for it.
Thanks to Jonathan Taylor for reproducing this bug.
Change-Id: I0fbbcc77d6945a38563d82efba4429f4b5278cd5
* Use relative error instead of absolute error.
* Update tolerance to account for embedded GPUs such as the Jetson TX2.
Change-Id: I05742ecbfd915e797fcbc4137acc5d1b513cd465
1. Add threading to all three subclasses of BlockRandomAccessMatrix.
i.e. BlockRandomAccessDenseMatrix, BlockRandomAccessSparseMatrix
and BlockRandomAccessDenseMatrix.
For BlockRandomAccessDenseMatrix and BlockRandomAccessSparseMatrix
this just means SetZero is parallelized. Which by itself is no
big deal, but by doing so, the constructor for all three subclasses
become uniform.
BlockRandomAccessSparseMatrix::SymmetricRightMultiplyAndAccumulate
maybe threaded in the future if needed.
BlockRandomAccessDiagonalMatrix is the biggest beneficiary. SetZero
Invert and RightMultiplyAndAccumulate are all threaded now.
2. Change the storage in BlockRandomAccessDiagonalMatrix from
TripletSparseMatrix to CompressedRowSparseMatrix. This has no
performance implications since we do not really use the capabilities
of the underlying matrix indexing representation. This is a forward
looking change when we decide to transfer this matrix to the GPU,
a CompressedRowSparseMatrix will save on a data conversion.
3. Use std::unique_ptr as needed and eliminate the need for custom
destructors.
4. Modify CompressedRowSparseMatrix::CreateBlockDiagonalMatrix to
take a nullptr as the data vector.
Fixes https://github.com/ceres-solver/ceres-solver/issues/936
Fixes https://github.com/ceres-solver/ceres-solver/issues/935
Change-Id: Ia6487f2d924fbe669835bdcc38abf2b451bda4ee
CoordinateDescentMinimizer optimizes one parameter block at a time.
To do this, it manipulates the parameter block object. It was doing
so inconsistently, where the tangent space offset was being set to
zero but the ambient state offset was not being set to zero. This
did not cause problems because these offsets were not really being
used inside the CoordinateDescentMinimizer. However the recent
change which parallelizes Program::Plus uncovered this bug.
The reason this bug was not caught was because, CoordinateDescentMinimizer
does not have any tests. I will fix this shortly, but in the interim
to unbreak inner iterations at head, this small change should go in.
Change-Id: I55d2698e8509f9cb5751e7a5180427129d86e720
Main focus of this change is to parallelize remaining operations (most of them
are operations on vectors) in code-path utilized with iterative Schur
complement.
Parallelization is handled using lazy evaluation of Eigen expressions.
On linux pc with intel 8176 processor parallelization of vector operations has
the following effect:
Running ./bin/parallel_vector_operations_benchmark
Run on (112 X 3200.32 MHz CPU s)
CPU Caches:
L1 Data 32 KiB (x56)
L1 Instruction 32 KiB (x56)
L2 Unified 1024 KiB (x56)
L3 Unified 39424 KiB (x2)
Load Average: 3.30, 8.41, 11.82
-----------------------------------
Benchmark Time
-----------------------------------
SetZero 10009532 ns
SetZeroParallel/1 10024139 ns
...
SetZeroParallel/16 877606 ns
Negate 4978856 ns
NegateParallel/1 5145413 ns
...
NegateParallel/16 721823 ns
Assign 10731408 ns
AssignParallel/1 10749944 ns
...
AssignParallel/16 1829381 ns
D2X 15214399 ns
D2XParallel/1 15623245 ns
...
D2XParallel/16 2687060 ns
DivideSqrt 8220050 ns
DivideSqrtParallel/1 9088467 ns
...
DivideSqrtParallel/16 905569 ns
Clamp 3502010 ns
ClampParallel/1 4507897 ns
...
ClampParallel/16 759576 ns
Norm 4426782 ns
NormParallel/1 4442805 ns
...
NormParallel/16 430290 ns
Dot 9023276 ns
DotParallel/1 9031304 ns
...
DotParallel/16 1157267 ns
Axpby 14608289 ns
AxpbyParallel/1 14570825 ns
...
AxpbyParallel/16 2672220 ns
-----------------------------------
Multi-threading of vector operations in ISC and program evaluation results into
the following improvement:
Running ./bin/evaluation_benchmark
--------------------------------------------------------------------------------------
Benchmark this 2fd81de
--------------------------------------------------------------------------------------
Residuals<problem-13682-4456117-pre.txt>/1 4136 ms 4292 ms
Residuals<problem-13682-4456117-pre.txt>/2 2919 ms 2670 ms
Residuals<problem-13682-4456117-pre.txt>/4 2065 ms 2198 ms
Residuals<problem-13682-4456117-pre.txt>/8 1458 ms 1609 ms
Residuals<problem-13682-4456117-pre.txt>/16 1152 ms 1227 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/1 19759 ms 20084 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/2 10921 ms 10977 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/4 6220 ms 6941 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/8 3490 ms 4398 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/16 2277 ms 3172 ms
Plus<problem-13682-4456117-pre.txt>/1 339 ms 322 ms
Plus<problem-13682-4456117-pre.txt>/2 220 ms
Plus<problem-13682-4456117-pre.txt>/4 128 ms
Plus<problem-13682-4456117-pre.txt>/8 78.0 ms
Plus<problem-13682-4456117-pre.txt>/16 49.8 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/1 2434 ms 2478 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/2 2706 ms 2688 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/4 1430 ms 1548 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/8 742 ms 883 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/16 438 ms 555 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/1 2438 ms 2481 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/2 2565 ms 2790 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/4 1434 ms 1551 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/8 765 ms 892 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/16 435 ms 559 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/1 1278 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/2 1555 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/4 833 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/8 459 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/16 250 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/1 1468 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/2 1871 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/4 957 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/8 528 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/16 294 ms
End-to-end improvements with bundle_adjuster invoked with
./bin/bundle_adjuster --num_threads 28 --num_iterations 40 \
--linear_solver iterative_schur \
--preconditioner jacobi --input
---------------------------------------------
Problem this 2fd81de
---------------------------------------------
problem-13682-4456117-pre.txt 508.6 892.7
problem-1778-993923-pre.txt 763.8 1129.9
problem-1723-156502-pre.txt 6.3 14.4
problem-356-226730-pre.txt 76.3 116.2
problem-257-65132-pre.txt 38.6 52.0
Change-Id: Ie31cc5015f13fa479c16ffb5ce48c9b880990d49