Commit Graph

1898 Commits

Author SHA1 Message Date
Joydeep Biswas c9d2ec8a9f Updates to sparse block matrix structures to support new
sparse linear solvers.

* Add methods to convert TripletSparseMatrix and BlockSparseMatrix to
  CRSMatrix structure.
* Added tests for conversion of TripletSparseMatrix and BlockSparseMatrix
  to CRSMatrix structure.
* Added documentation on the BlockSparseMatrix structure.

Change-Id: I020cfa91c301567ceeb39ff2064183c5d88c9ed5
2022-08-06 22:03:37 -05:00
Sergiu Deitsch 5fe0bd45a9 Added MinGW to Windows Github workflow
Change-Id: Id2bcd92a5464ac4888295c3dbfe9b806be95a3d9
2022-08-06 23:56:51 +02:00
Sameer Agarwal 738c027c1f Fix a logic error in iterative_refiner_test
Change-Id: I741802db37d6d9e42e38cef358e00369f6a38d06
2022-08-06 08:29:08 -07:00
Sameer Agarwal cb6ad463d0 Add mixed precision support for CPU based DenseCholesky
On problem-744-543562-pre.txt

The time spent in linear solver on my M1 Pro is

eigen        81.550970
eigen+mixed  54.107383
LAPACK       47.078127
LAPACK+mixed 28.639868

Solution quality is unaffected.

The implementation of RefinedDenseCholesky and DenseIterativeRefiner
are straightforward ports of RefinedSparseCholesky and
SparseIterativeRefiner (formerly IterativeRefiner).

It maybe possible to refactor the SparseCholesky and DenseCholesky
interfaces so that this code duplication can be removed in the
future.

Change-Id: I921334224cb97629a60390f2add822de207f7923
2022-08-05 15:30:12 -07:00
Julio L. Paneque df55682ba5 Fix Eigen error in 2D sphere manifolds
Since Eigen does not allow to have a RowMajor column vector (see
https://gitlab.com/libeigen/eigen/-/issues/416), the storage order
must be set to ColMajor in that case. This fix adds that special
case when generating 2D sphere manifolds.

Change-Id: I594932e0dafc878e0b348f72524478588e61b34d
2022-08-05 11:25:29 +02:00
Sergiu Deitsch 1cf59f61eb Set Github workflow NDK path explicitly
Change-Id: I31115293f4a80ca15a7520814055ee1268d2970f
2022-08-01 18:27:08 +02:00
Sameer Agarwal 68c53bb395 Remove ceres::LocalParameterization
Change-Id: I3bdf2f6a8857db10c984024a27f490eefd23fefa
2022-07-29 22:29:23 +00:00
Joydeep Biswas 2f660464cc Fix build issue with CUDA testing targets when compiling without gflags.
Change-Id: I926a85c30b51802a99161679cb2a28fda6e3ef47
2022-07-20 21:31:47 +05:30
Sameer Agarwal c801192d47 Minor fixes
Change-Id: I4c825bbd19b2d902d17dce37d228e23a808c87fb
2022-07-18 06:43:12 -07:00
Sameer Agarwal ce9e902b86 Fix missing CERES_METIS_VERSION
CERES_METIS_VERSION needs to be set if either Eigen or
SuiteSparse are using it. Previously, we were conditioning it
only on EIGENMETIS being ON.

Change-Id: I380a3b138b79903aa142b560ed46f438ff549a82
2022-07-14 14:21:02 +00:00
Sameer Agarwal d9a3dfbf20 Add a missing ifdef guard to dense_cholesky_test
Change-Id: Ibd924bc52b589ea70c5e7a543e1092697a0e5941
2022-07-14 06:58:49 -07:00
Sameer Agarwal 5bd43a1fa4 Speed up DenseSparseMatrix::SquareColumnNorm.
Because we store the matrix as row major matrix, the
obvious Eigen expression performs rather poorly. A straight
c++ loop speeds things up considerably.

Also replace use of matrix() with direct use of m_.

Change-Id: I3d6166df4765ad8400ab9602a54b65fd21b1d50f
2022-07-14 06:46:49 -07:00
Sameer Agarwal cbc86f6512 Fix the build when CUDA is not present
Change-Id: Ieaa483ce190ba917096c675f7ee731cb20d27bf2
2022-07-13 10:01:03 -07:00
Sameer Agarwal 5af8e64497 Update year in solver.h
Change-Id: I49ffe5a24786aa32e7677785cc91ee914956f2e6
2022-07-13 09:09:37 -07:00
Joydeep Biswas 88e08cfe71 Mixed-precision Iterative Refinement Cholesky With CUDA
* Created a new class CUDADenseCholeskyMixedPrecision, which performs
  Cholesky factorization and solving in single (fp32) precision, and
  optionally performs iterative refinement.
* Added CUDA kernels for mixed-precision solve operations
* Added more detailed timing information to the FullReport about Schur
  elimination, reduced system solves, and back-substitution.

Some test performance numbers follow.
All tests were performed on an Ubuntu 20.04 desktop with an
Intel Core i9-9940X CPU and Nvidia Quadro RTX 6000 GPU.

Tests were launched as:
./bin/bundle_adjuster --input (problem_file) \
    --num_iterations 20
    --num_threads 28
    --linear_solver dense_schur
    --dense_linear_algebra_library (cuda|lapack)
    [--mixed_precision_solves]

==================================================
problem-21-11315-pre.txt
==================================================

--------------------------------------------------
Cuda Mixed Precision
--------------------------------------------------
Cost:
Initial                          4.413239e+06
Final                            3.037864e+04
Change                           4.382861e+06
  Linear solver                      0.250703 (14)
  ├ Schur eliminate                  0.234025 (14)
  ├ Reduced solve                    0.006643 (14)
  └ Backsubstitute                   0.006598 (12)

--------------------------------------------------
Cuda
--------------------------------------------------
Cost:
Initial                          4.413239e+06
Final                            3.037864e+04
Change                           4.382861e+06
  Linear solver                      0.257517 (12)
  ├ Schur eliminate                  0.233518 (12)
  ├ Reduced solve                    0.010621 (12)
  └ Backsubstitute                   0.007124 (12)

--------------------------------------------------
Lapack (OpenBLAS)
--------------------------------------------------
Cost:
Initial                          4.413239e+06
Final                            3.037864e+04
Change                           4.382861e+06
  Linear solver                      0.332349 (12)
  ├ Schur eliminate                  0.274748 (12)
  ├ Reduced solve                    0.015966 (12)
  └ Backsubstitute                   0.034192 (12)

==================================================
problem-257-65132-pre.txt
==================================================

--------------------------------------------------
Cuda Mixed Precision
--------------------------------------------------
Cost:
Initial                          2.456242e+07
Final                            9.677593e+04
Change                           2.446565e+07
  Linear solver                      1.332367 (20)
  ├ Schur eliminate                  1.021365 (20)
  ├ Reduced solve                    0.195472 (20)
  └ Backsubstitute                   0.075582 (20)

--------------------------------------------------
Cuda
--------------------------------------------------
Cost:
Initial                          2.456242e+07
Final                            9.677547e+04
Change                           2.446565e+07
  Linear solver                      1.810176 (20)
  ├ Schur eliminate                  1.012862 (20)
  ├ Reduced solve                    0.678704 (20)
  └ Backsubstitute                   0.083925 (20)

--------------------------------------------------
Lapack (OpenBLAS)
--------------------------------------------------
Cost:
Initial                          2.456242e+07
Final                            9.677547e+04
Change                           2.446565e+07
  Linear solver                      2.376273 (20)
  ├ Schur eliminate                  0.987613 (20)
  ├ Reduced solve                    1.043873 (20)
  └ Backsubstitute                   0.310402 (20)

==================================================
problem-744-543562-pre.txt
==================================================

--------------------------------------------------
Cuda Mixed Precision
--------------------------------------------------
Cost:
Initial                          1.434881e+08
Final                            1.546895e+06
Change                           1.419412e+08
  Linear solver                     27.010088 (20)
  ├ Schur eliminate                 24.362433 (20)
  ├ Reduced solve                    1.428542 (20)
  └ Backsubstitute                   0.814266 (20)

--------------------------------------------------
Cuda
--------------------------------------------------
Cost:
Initial                          1.434881e+08
Final                            1.546895e+06
Change                           1.419412e+08
  Linear solver                     32.342513 (20)
  ├ Schur eliminate                 24.638819 (20)
  ├ Reduced solve                    6.492090 (20)
  └ Backsubstitute                   0.802184 (20)

--------------------------------------------------
Lapack (OpenBLAS)
--------------------------------------------------
Cost:
Initial                          1.434881e+08
Final                            1.546895e+06
Change                           1.419412e+08
  Linear solver                     34.152224 (20)
  ├ Schur eliminate                 24.183723 (20)
  ├ Reduced solve                    8.784413 (20)
  └ Backsubstitute                   0.795044 (20)

Change-Id: I178887e776d8f4a1e8abb99bbc205bf8c278bf79
2022-07-13 06:55:31 -05:00
Alex Stewart 290b34ef05 Fix optional SuiteSparse + METIS test-suite names to be unique
Change-Id: I539dc3edebbf1929712d5187e66ebf9f615844f8
2022-07-08 16:55:36 +01:00
Alex Stewart d038e2d837 Fix use of NESDIS with SuiteSparse in tests if METIS is not found
Change-Id: I6da004d091a463485935b7f7fa45e56dfcd4341c
2022-07-08 15:26:10 +01:00
Sergiu Deitsch 027e741a1a Eliminated MinGW warning
Change-Id: I35a852e742cc7d678c3af40e1bcd8a4f962303ee
2022-06-26 05:09:09 +00:00
Sergiu Deitsch 4e5ea292ba Fixed MSVC 2022 warning
MSVC rightfully issues warning C4305: 'if': truncation from 'size_t' to
'bool' in a static_assert condition that implicitly converts sizeof
result to a boolean.

Change-Id: Ie3b913288bfeaa7a4b362ef7f83d2505ed368641
2022-06-24 00:12:09 +02:00
Alex Stewart 83f6e08530 Fix use of conditional preprocessor checks within a macro in tests
- These are non-standard C++, and whilst they are accepted by GCC and
  Clang on *NIX and macOS, they are rejected by MSVC.

Change-Id: Ie627d74bb02ebdce3dc5e13c2010616c26cb5dea
2022-06-23 16:30:59 +01:00
Alex Stewart 70f1aac31f Fix fmin/fmax() when using Jets with float as their scalar type
Change-Id: Ie7f400763f91b0e264a50401716321587ab1d477
2022-06-23 16:05:58 +01:00
Alex Stewart 5de77f399e Fix reporting of METIS version
- Also fixes behaviour of EIGENMETIS option to match that of the other
  CMake dependency options, and ensure that its value aligns exactly
  with whether Eigen support for METIS will be compiled into Ceres.

Change-Id: Ifbf6f5d82b9ba89a156673eb6042519a985e6b04
2022-06-22 19:20:40 +01:00
Alex Stewart 11e6376675 Fix #ifdef guards around METIS usage in EigenSparse backend and tests
Change-Id: Idd1e0b7b1b2df2d402431b107dee0fa8e0fd58f5
2022-06-22 19:05:30 +01:00
Sergiu Deitsch 0c88301e66 Provide optional METIS support
* Split `CERES_NO_METIS` into two defines: `CERES_NO_PARTITION` and
  `CERES_NO_METIS`. The former refers to METIS support in SuiteSparse,
  the latter to the Eigen's MetisSupport module. This enables the use of
  sparse matrix reordering independent from SuiteSparse.
* Run Linux, macOS, and macOS Github workflows with METIS enabled
  SuiteSparse.

Fixes #808

Change-Id: I5076b7e1268d32cc3e7e56650edcbaf7fb3b59ce
2022-06-22 16:46:02 +00:00
Alex Stewart f11c256265 Fix fmin/fmax() to use Jet averaging on equality
- Prior to 48cb54d1, Ceres' fmin/fmax() for Jets followed the convention
  of std::min/max(), and always returned the first argument on equality,
  irrespective of whether this argument was natively a scalar or a Jet.
- After 48cb54d1, Ceres' fmin/fmax() instead returned the second
  argument on equality, again irrespective of whether this argument was
  natively a scalar or a Jet.
- Now on equality we average the arguments as Jets, which ensures that
  a consistent answer is produced irrespective of the ordering or type
  (Jet or scalar) of the input arguments. This also ensures that we
  preserve a non-zero derivative where it exists, excluding the edge
  case of two Jet inputs with equal but oppositely signed infinitesimal
  components.
- We retain the behaviour introduced in 48cb54d1 whereby NaNs are
  treated as missing values, following the convention of
  std::fmin/fmax().
- Raised as issue #816.

Change-Id: I01217c0e32c1be83be440e4515b57c79dd290923
2022-06-22 14:19:55 +01:00
Joydeep Biswas b90053f1ad Revert C++17 usage of std::exclusive_scan
* Unfortunately on some systems such as the Nvidia Jetson, while the
  compiler supports C++17, the STL implementations are incomplete.
  One such missing implementation is std::exclusive_scan, so this
  patch reverts to the old way of manually computing prefix sums.

Change-Id: I4192257519b0083560a4b44e2659ee44d7421105
2022-06-12 11:24:08 -05:00
Sergiu Deitsch dfce1e128d Link against threading library only if necessary
1. The platform specific threads library is only needed if we actually
   use threads. In this case, the library is not optional opposed to
   previous logic.
2. Do not hide the find module output to allow the user to understand
   what happens in case of a CMake failure to locate Threads.
3. Finally, Threads is private dependency that does need to be
   propagated to consumers unless Ceres was compiled as a static
   library.

Change-Id: I8d9d9cd42930e1ed234f69a2dba70d0ee2755b4e
2022-06-08 00:03:41 +02:00
Sergiu Deitsch 69eddfb6da Use find module to link against OpenMP
Depending on the compiler in use, linking against OpenMP may require
passing specific compiler flags instead of linking against a library.
Use the CMake OpenMP find module to abstract OpenMP activation.

Change-Id: Ib43f576ac12e2c5e9598e9586df3dfa018e9c08b
2022-06-07 23:39:39 +02:00
Sameer Agarwal b4803778c3 Update documentation for linear_solver_ordering_type
Also update obsolete documentation related to building and
using sparse linear algebra libraries.

Change-Id: I83682b43472e6a6ec4e4dad32fa21c089d518c06
2022-06-07 14:08:26 -07:00
Joydeep Biswas 2e764df06f Update Cuda memcheck test
* Fix silly typo in CMakeLists.txt

Change-Id: I98b5a2fc0b8452f2f078117e31fb0ef350e6c11f
2022-06-02 17:41:40 -05:00
Joydeep Biswas 443ae9ce26 Update Cuda memcheck test
* Previously the Cuda memcheck tests relied on the Cuda binaries being
  on the environment PATH. This has been changed instead to use the
  path discovered by CMake when searching for Cuda. This has the added
  benefit that the memcheck tool will be sure to be from the same Cuda
  version install as the version being compiled against.

Change-Id: I650d1bb7e14064ca98a01e3c13eb1bcb772b51cc
2022-06-02 22:25:58 +00:00
Sergiu Deitsch 55b4c3f447 Retain terminal formatting when building docs
This prevents Sphinx output to be stripped of colors, emphasis etc.

Change-Id: I127e02cbdda69a5d49a73678a7a7e2b4512b189e
2022-05-28 21:02:12 +00:00
Sergiu Deitsch 786866d9f7 Generate version string at compile time
Strings can be concatenated during compilation bypassing any dynamic
memory allocation.

Change-Id: Iecd94ca44dddde4694bfeb823a0a06f174b6085b
2022-05-28 15:04:09 +02:00
Sameer Agarwal 5bd83c4ac0 Unbreak the build with EIGENSPARSE is disabled
Change-Id: Ia3a6121f031e647b51adba427c814e818eed2d2d
2022-05-27 10:12:50 -07:00
Sameer Agarwal 2335b5b4b7 Remove support for CXSparse
Eigen provides all the functionality that we need from CXSparse
with a more liberal license.

I will update the documentation in a follow up CL.

Change-Id: I0b9fd8be3c27754cc2986cc0e06595c8b3fdec0b
2022-05-27 09:20:21 -07:00
Sameer Agarwal fbc2eea166 Nested dissection for ACCELERATE_SPARSE & EIGEN_SPARSE
Change-Id: Iec8ea6b0a537559b48b59bcfc91b94b58cb2070e
2022-05-27 06:50:12 -07:00
Sergiu Deitsch d87fd551bc Fix Ubuntu 20.04 workflow tests
Previously, the tests did not run because the CMake version shipped with
Ubuntu 20.04 does not understand the `--test-dir` option and silently
fails.

Change-Id: I335e1d9e3890aa56e66a9dfd0fccd4594a84a08c
2022-05-27 10:47:07 +00:00
Sergiu Deitsch 71717f37c6 Use glog 0.6 release to run Windows Github workflow
Change-Id: I2853ea13c798c6ef05b3a19a2d3986b514caf1ed
2022-05-27 11:43:41 +02:00
Sameer Agarwal 66e0adfa70 Fix detection of sphinx-rtd-theme
Upstreaming fix from Debian.

https: //github.com/ceres-solver/ceres-solver/issues/809
Change-Id: I0e2f90a405a56ceffdda37f70d6e1ac853e176f1
2022-05-23 23:02:40 -07:00
Sameer Agarwal d09f7e9d5e Enable postordering when computing the sparse factorization.
Previously when using a natural ordering, we had postordering
turned off. This is not a good idea. Enabling postordering will
also has the possibility of improving the size of the supernodes.

Change-Id: I8c270e54751b8bed53b38a0b461f647f5c8f5640
2022-05-21 14:23:36 -07:00
Sameer Agarwal 9b34ecef1c Unbreak the build on MacOS
Change-Id: I9144a84842baf1921b8d5808983d8d4e7cde747a
2022-05-19 14:21:15 -07:00
Sameer Agarwal 8ba8fbb173 Remove Solver::Options::use_postordering
This was an ill-advised and complicated to interpret option
which offers nothing particularly useful.

Change-Id: Ia7741ed62ef977c96fa52299a884e404bee659ac
2022-05-19 21:10:33 +00:00
Sameer Agarwal 30b4d5df35 Fix the ceres.bzl to add missing cc files.
Thanks to nate-thirdwave@ for pointing this out and offering
a fix.

Also add a TODO about an odd loop in covariance_impl.cc which was
revealed as I was testing the bazel build

https: //github.com/ceres-solver/ceres-solver/issues/800
Change-Id: I87d17155ee43ea2a52b8031177d6b3ac5ae1460a
2022-05-19 14:07:15 -07:00
Sameer Agarwal 39ec5e8f99 Add Nested Dissection based fill reducing ordering
With this change, the user can now choose between Approximate Minimum
Degree and Nested Dissection as a fill reducing algorithm when using
a sparse direct factorization based linear solver like SPARSE_NORMAL_CHOLESKY
or SPARSE_SCHUR.

Currenly only SUITE_SPARSE is supported. It requires that
SuiteSparse be compiled with Metis support enabled.

On most problems AMD is still the better choice, but in some cases
like the grid3D dataset from https://lucacarlone.mit.edu/datasets/
the solution time with AMD is 57s and with NESDIS 38 on my M1 Mac.

On some other problems at Google we have observed speedups of 10x,
there is also a corresponding decrease in the total amount of memory
used.

This patch is based on the original work done by NeroBurner in
https://ceres-solver-review.googlesource.com/c/ceres-solver/+/20580

1. Add a new enum to the public api LinearSolverOrderingType and
   a setting Solver::Options::linear_solver_ordering_type.
2. TrustRegionPreprocessor had some complicated logic which determined
   when linear solvers should reorder their matrices on their own and not
   this has been refactored into a more readable function that lives
   inside reorder_program.h/cc.
3. Plumbing in reorder_program.cc and trust_region_processor.cc to use
   nested dissection.
4. Update bundle_adjuster.cc to use nested dissection.

Change-Id: I388b027934f86c58b4da2b65a4fa5204ea73bf40
2022-05-19 12:36:20 -07:00
Sameer Agarwal aa62dd86a8 Fix a build breakage
Change-Id: I57591bc42d53f9856b49f7a16732a2a1e259dc67
2022-05-19 11:20:38 -07:00
Sameer Agarwal 41c5fb1e80 Refactor suitesparse.h/cc
1. Generalize SuiteSparse::AnalyzeCholesky and
   SuiteSparse::BlockAnalyzeCholesky from just doing AMD to taking
   OrderingType as an argument and using that to determine whether
   AMD & Nested Dissection algorithms are used for computing the
   fill-reducing ordering or a natural ordering when computing
   the symbolic factorization.

2. Remove AnalyzeCholeskyWithNaturalOrdering.

3. Replace and generalize SuiteSparse::BlockAMDOrdering with
   SuiteSparse::BlockOrdering which also takes OrderingType as an
   argument. Same for SuiteSparse::ApproximateMinimumDegreeOrdering
   and SuiteSparse::NestedDissectionOrdering by
   SuiteSparse::Ordering.

4. Remove LinearSolver::Options::use_postordering and replace it
   with LinearSolver::Options::ordering_type.

5. Replace Preconditioner::Options::use_postordering and replace it
   with Preconditioner::Options::ordering_type.

6. Add NESDIS to OrderingType. With the above changes, the linear
   solvers can now use Nested Dissection once this information
   is piped through the nonlinear solver.

Change-Id: Ib8e93fbf34ae2981bf2ac54dcda9e25c7c213790
2022-05-19 11:05:46 -07:00
Sameer Agarwal 12263e2830 Make the min. required version of SuiteSparse to be 4.5.6
With this change we can drop the complicated/conditional handling
around CAMD and assume that it is always available.

Change-Id: I93e1da676fb75817f79824b8b2b6549d03f278b0
2022-05-16 12:48:43 -07:00
Sameer Agarwal c8493fc366 Convert internal enums to be class enums.
Change-Id: Ide89c7115c3b12c0f2452a2969dc5523b3a7970f
2022-05-16 12:47:15 -07:00
Sameer Agarwal bb3a40c091 Add Nested Dissection ordering method to SuiteSparse
Change-Id: I5e00977839d9d5ce914bda0978d81e97e28fc673
2022-05-14 15:08:10 -07:00
Evan Levine f1414cb5bd Correct spelling in comments and docs.
Change-Id: Iad9a0599d644d3b3cd54244edaf64d408cb1308e
2022-04-24 21:40:13 -07:00