Commit Graph

2080 Commits

Author SHA1 Message Date
Sameer Agarwal bea2477010 Add a missing include dir to the cuda kernels target.
This was somehow overlooked in the last patch causing the cuda
built to break.

Change-Id: Ic447103b139f6a7b607d664d02b9cbae98072ace
2023-09-22 15:28:12 -07:00
Dmitriy Korchemkin 18ea7d1c21 Runtime check for cudaMallocAsync support
Change-Id: Ia0e347d99b005d805ff2351cdb8918cc1331fc24
2023-09-22 19:25:25 +00:00
Sameer Agarwal a227045be1 Remove cuda-memcheck based tests
cuda-memcheck has been deprecated and these tests will be reinstated
once we move to compute sanitizer.

Change-Id: I7e01cffdd00ffb7404dfef2a171ebaeca017cda8
2023-09-22 06:09:04 -07:00
Sameer Agarwal d10e786ca8 Remove an unused variable from CudaSparseMatrix
Change-Id: Id9c969fda8d636dc793b98b452b2aaffacb11180
2023-09-21 22:33:38 -07:00
Sameer Agarwal 5a30cae583 Preparing for 2.2.0rc1
1. Add a version history
2. Update copyright years across the code base
3. Run format_all.sh
4. Update version strings from 2.1.0 to 2.2.0 in the docs and
   elsewhere.

Change-Id: I46d8d479d54bd6002d532785e67342106e73c9ac
2.2.0rc1
2023-09-21 11:23:38 -07:00
Mark Shachkov 9cca671273 Enable compatibility with SuiteSparse 7.2.0
Change-Id: I072dc3f7c245fc2ebbdffed715ac4def20f7dccd
2023-09-17 20:57:43 +02:00
Sergiu Deitsch a1c02e8d37 Rework the Sphinx find module
* Make sphinx_rtd_theme a find module component to avoid hard-wiring it
  into the module and allowing to report the theme in case it is missing
  using the standard CMake package mechanism.
* Adjust find module cache variables names case to match the find module
  name.
* Also report sphinx-build version for completeness.
* Invoke the find module only once. Calling find_package on the same
  module is not needed.

Change-Id: I9d1bf0fcc0d44b9b37e624128812f348c5442ada
2023-09-12 20:51:48 +02:00
Sergiu Deitsch a57e35bbab Require at least CMake 3.16
Given we no longer support Ubuntu 18.04 due to packaged GCC lacking
C++17 support we can bump the minimum required CMake version to the one
provided by Ubuntu 20.04 which is CMake 3.16. Consequently, this allows
to drop some of the legacy CMake logic.

Change-Id: I1f05d4c5681d10aa7faa0800ef4a803be2f5b7dd
2023-09-12 19:33:00 +02:00
Sergiu Deitsch 863db948f3 Eliminate macOS sprintf warning
AppleClang 14.0.0.14000029 warns about a potential security problem
while invoking the sprintf C function:

    internal/ceres/fixed_array_test.cc:469:3: warning: 'sprintf' is deprecated: This function is provided for compatibility reasons only.  Due to security concerns inherent in the design of sprintf(3), it is highly recommended that you use snprintf(3) instead. [-Wdeprecated-declarations]
      sprintf(buf.data(), "foo");  // NOLINT(runtime/printf)
      ^
    /Applications/Xcode_14.2.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX13.1.sdk/usr/include/stdio.h:188:1: note: 'sprintf' has been explicitly marked deprecated here
    __deprecated_msg("This function is provided for compatibility reasons only.  Due to security concerns inherent in the design of sprintf(3), it is highly recommended that you use snprintf(3) instead.")
    ^
    /Applications/Xcode_14.2.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX13.1.sdk/usr/include/sys/cdefs.h:215:48: note: expanded from macro '__deprecated_msg'
            #define __deprecated_msg(_msg) __attribute__((__deprecated__(_msg)))

Replace sprintf by snprintf to avoid this deprecation warning.

Change-Id: I6870c0bd4e390388d1d7bcec082cee272b234eba
2023-09-11 22:51:28 +02:00
Sergiu Deitsch 6f5342db68 Export ceres_cuda_kernels to project's build tree
This avoids the following CMake error when exporting Ceres build
directory to local CMake package registry:

    export called with target "ceres" which requires target
      "ceres_cuda_kernels" that is not in any export set.

Fixes #966

Change-Id: I5a75191fc414a3f138b19cb9d5850f7330f2d24a
2023-09-11 00:17:04 +02:00
Sergiu Deitsch d864d146fd Add macOS 13 runner to Github workflows
Change-Id: Ied8e255bb5d8fbeaa0c43c6af22775dda8aad1e8
2023-09-10 21:05:03 +02:00
Sameer Agarwal 01a23504d5 Add a workaround for CMake Policy CMP0148
Without this we start getting errors related to FindSphinx.cmake

I am not sure yet, what version of cmake we can assume in the wild
but the current minimum version supports the old behaviour and this
policy seems relatively recent (CMake version 3.27) so we should
have this workaround till we update our minimum required version.

Fixes https://github.com/ceres-solver/ceres-solver/issues/1002

Change-Id: I1beaac9ee27bc9ff85b64f53f72606a0424f2391
2023-09-09 07:58:49 -07:00
Sergiu Deitsch de9cbde95d Work around MinGW32 manifold_test segfault
Converting fixed size vectors to dynamic ones allows to avoid
segmentation faults in Eigen's packet math if the corresponding
expressions are invoked within GMock matchers.

Fixes #996

Change-Id: I7da5599883825ab0e580678d3d55de19095b41b1
2023-09-08 19:57:43 +02:00
Dmitriy Korchemkin 5e4b22f7fc Update CudaSparseMatrix class
- Perform temporary buffer size estimation only once
- Allow construction from existing buffers with col/row structure

Change-Id: I73c291328f1e8ed9184aba5d7058df71cbc6a15d
2023-08-31 18:56:44 +00:00
MaximSmolskiy ed9921fc24 Fix Solver::Options in documentation
Change-Id: Ia01fdba7561aac539ce96f09b3d404cbe061de11
2023-08-29 05:53:36 +03:00
Sameer Agarwal de62bf2204 Two minor fixes
1. In cuda_sparse_matrix.cc fix the order of fields in the initializer list.
2. Move a line of code to the ifdef branch which will use it.

Change-Id: If32ea14a287f845c1740e6f726c1007e86a4eeca
2023-08-20 17:53:10 +00:00
MaximSmolskiy 5f97455bea Fix typos in documentation
Change-Id: I1d39fe777c87d3480bc403ec4602dd2dd9284839
2023-08-20 20:42:05 +03:00
Dmitriy Korchemkin 5ba62abecb Add CUDA support to windows CI builds
Change-Id: I0fcadfbc1eef39b8aeaee119f2b42bc8f4a74314
2023-08-16 17:54:15 +00:00
Jason Mak a98fdf5822 Update CMakeLists.txt to fix Windows CUDA build
- In the top-level CMakeLists.txt, certain flags are passed to
  disable warnings and increase the maximum size of an object file.
  nvcc cannot handle these flags so we tell CMake to only use them
  for C++ code (and not CUDA code).
- In internal/ceres/CMakeLists.txt, CMake is originally told to link
  the import library cudart.lib when linking CUDA code. By default,
  it seems that Visual Studio will link the static library
  cudart_static.lib when linking CUDA code. So we avoid linking with
  cudart.lib to avoid linking the same library twice.

Change-Id: I1fbf0d7e76d57b4338708757b27f5074722608cb
2023-08-16 17:10:09 +00:00
Dmitriy Korchemkin ec4907399a Fix block-sparse to crs conversion on windows
Change-Id: I2aabdb68afc8152c5d8da157baf22faf8d8f80cf
2023-08-16 16:16:41 +00:00
Dmitriy Korchemkin 799ee91bbf Fix check in CompressedRowJacobianWriter::CreateJacobian()
Change-Id: I8c8d418f14a9e489a94b5117460e739f65e36a96
2023-08-16 11:14:42 +00:00
Sameer Agarwal ee90f6cbf7 Detect large Jacobians and return failure instead of crashing.
Detect when the number of non-zeros overflows when constructing
BlockSparseMatrix and CompressedRowSparseMatrix and return
with an error message instead of crashing.

Change-Id: I45e102f7c0519eef441ce0586b7adf96e4a954a9
2023-08-11 17:16:28 -07:00
Sameer Agarwal 310a252fb6 Deal with infinite initial cost correctly.
Previously it could be the case that a residual block could return
a residual whose squared norm overflows and generates an infinity
which we did not detect. This would then lead to the trust region
minimizer incorrectly terminating indicating convergence while
generating a cost delta of NaN.

This change adds a check for that and also does two minor cosmetic
changes.

1. Reduce the level of nesting in program_evaluator.h by adding
   an early return.
2. The error message when IterationZero fails now says that the
   Initial residual and Jacobian failed, to indicate that the
   optimizer had no chance to do any work.

Fixes https://github.com/ceres-solver/ceres-solver/issues/988

Thanks to @Ashray-g for reporting this.

Change-Id: I52ae7627a66f637135209dbb2e42935b52c8bc77
2023-08-04 18:35:25 -07:00
Sameer Agarwal 357482db70 Add NumTraits::max_digits10 for Jets
Change-Id: Ief2c02eeaefba98d2540eb2f3428c65a98e36ef5
2023-07-29 12:35:30 -07:00
MaximSmolskiy 908b5b1b53 Fix type mismatch in documentation
Change-Id: Ief4d6a5039ea7af6f0b318df5146abe804611ca3
2023-07-16 00:49:42 +03:00
Dmitriy Korchemkin 75bacedf7d CUDA partitioned matrix view
Converts BlockSparseMatrix into two instances of CudaSparseMatrix,
corresponding to left and right sub-matrix.

Values of submatrix E are always just copied as-is, and values of
submatrix F are copied if each row-block of F submatrix satisfies
at least one of the following conditions:
 - There is atmost one cell in row-block
 - Row block has height of 1 row
Otherwise, indices of values in CRS order corresponding to value indices
in block-sparse order are computed on-the-fly.

Change-Id: I14eee00c36ee74b6b83fc85927907641383abfc7
2023-07-14 18:12:20 +00:00
Sergiu Deitsch fd6197ce0e Fixed GCC 13.1 compilation errors
Change-Id: I4c0ad38c899a9e5988e8e9f7cb870116aa8233eb
2023-06-10 17:37:51 +02:00
Dmitriy Korchemkin df97a8f059 Improve support of older CUDA toolkit versions
- Provides fall-back for older versions of CUDA toolkit
 - Using older versions of CUDA toolkit might result in
   over-synchronization

Change-Id: I545e6625d2342be30cb759b90bda379e555d7370
2023-05-31 22:45:35 +03:00
Dmitriy Korchemkin 085214ea74 Fix test CompressedRowSparseMatrix.Transpose
Change-Id: I653b4619d518f767e57f04136117d7fe47f33ca1
2023-05-30 16:03:57 +03:00
Dmitriy Korchemkin d880df09f9 Match new[] with delete[] in BSM
Change-Id: If78911c9570ce6a9039192501e6da7db3974293a
2023-05-30 11:28:33 +03:00
Dmitriy Korchemkin bdee4d6172 Block-sparse to CRS conversion using block-structure
Instead of pre-computing pemutation from block-sparse to CRS order,
index of value in CRS matrix is computed in the process of updating
values using block-sparse structure.

When it is possible to update values via a simple host-to-device copy,
block-sparse structure on GPU is discarded after computing CRS
structure.

Computing index is significantly slower than using pre-computed
permutation, but is still hidden by host-to-device transfer.

On problems from BAL dataset this results into reduction of extra
gpu memory consumption from 33% (permutation stored as 32-bit indices)
to ~10% for storing block-sparse structure.

Benchmark results:

======================= CUDA Device Properties ======================
Cuda version         : 11.8
Device ID            : 0
Device name          : NVIDIA GeForce RTX 2080 Ti
Total GPU memory     :  11012 MiB
GPU memory available :  10852 MiB
Compute capability   : 7.5
Warp size            : 32
Max threads per block: 1024
Max threads per dim  : 1024 1024 64
Max grid size        : 2147483647 65535 65535
Multiprocessor count : 68
====================================================================
Running ./bin/evaluation_benchmark
Run on (112 X 3200 MHz CPU s)
CPU Caches:
  L1 Data 32 KiB (x56)
  L1 Instruction 32 KiB (x56)
  L2 Unified 1024 KiB (x56)
  L3 Unified 39424 KiB (x2)
Load Average: 24.58, 11.75, 8.52

-----------------------------------------------------------------------
Benchmark                                                          Time
-----------------------------------------------------------------------
Using on-the-fly computation of CRS index corresponding to block-sparse
index:

JacobianToCRS<g/final/problem-4585-1324582-pre.txt>             1607 ms
JacobianToCRSView<g/final/problem-4585-1324582-pre.txt>          564 ms
JacobianToCRSMatrix<g/final/problem-4585-1324582-pre.txt>       2226 ms
JacobianToCRSViewUpdate<g/final/problem-4585-1324582-pre.txt>    228 ms
JacobianToCRSMatrixUpdate<g/final/problem-4585-1324582-pre.txt>  400 ms

Using precomputed permutation:
JacobianToCRS</final/problem-4585-1324582-pre.txt>              1656 ms
JacobianToCRSView</final/problem-4585-1324582-pre.txt>           553 ms
JacobianToCRSMatrix</final/problem-4585-1324582-pre.txt>        2255 ms
JacobianToCRSViewUpdate</final/problem-4585-1324582-pre.txt>     228 ms
JacobianToCRSMatrixUpdate</final/problem-4585-1324582-pre.txt>   406 ms

Performance of JacobianToCRSViewUpdate is still limited by
host-to-device transfer, and JacobianToCRSView is faster than computing
CRS structure on CPU.

Change-Id: Ifb6910fb01ae6071400d36c277846fadc5857964
2023-05-26 01:12:47 +03:00
Sameer Agarwal 0f9de3daf4 Use page locked memory in BlockSparseMatrix
If using CUDA_SPARSE for an iterative solve on the GPU,
allocate the values array in BlockSparseMatrix to make copying
to the GPU faster.

Change-Id: I63c1d2512babd74fc275b277ac8c3eabf3ec1144
2023-05-15 12:25:38 -07:00
Dmitriy Korchemkin e7bd72d41e Permutation-based conversion from block-sparse to crs
Change-Id: Ic33a6476c033187dff61886deb6d1761524943f0
2023-05-12 03:33:25 +03:00
Dmitriy Korchemkin abbc4e7974 Explicitly compute number of non-zeros in row
Change-Id: Iaebaf6d23d33dbbbe5a7c7240319bf28cf2bdd3a
2023-05-04 21:16:39 +03:00
Hs293Go 96fdfd2e7a Implement tests for Euler conversion with jets
https://github.com/ceres-solver/ceres-solver/issues/965

Change-Id: I6bdabf5ee09c000e49ea1757fde724d4bc4ccb92
2023-04-27 16:42:25 -04:00
Sameer Agarwal db1ebd3ff5 Work around lack of constexpr constructors for Jet
https://github.com/ceres-solver/ceres-solver/issues/965

Change-Id: Ia74a64568815605586c1f326a55cc36b09f33f30
2023-04-18 16:26:24 -07:00
Sameer Agarwal 16a4fa04e2 Further Jet conversion fixes
https://github.com/ceres-solver/ceres-solver/issues/965

Change-Id: Ib80467eda4120b0483814e9e01e9644dd7ee91a7
2023-04-18 14:58:22 -07:00
Sameer Agarwal 92ad18b8ae Fix a Jet conversion bug in rotation.h
https: //github.com/ceres-solver/ceres-solver/issues/965
Change-Id: I4c99639be6afbfbefea3f9d56b6eaddbe4f23fe9
2023-04-18 09:18:01 -07:00
Sameer Agarwal a5e745d4e1 ClangTidy fixes
Change-Id: I43badfd6706a44148989ac5621134453bad28133
2023-04-18 06:37:52 -07:00
Dmitriy Korchemkin 77ad8bb4e5 Change storage in BlockRandomAccessSparseMatrix
- TripletSparseMatrix in BlockRandomAccessSparseMatrix is replaced with
   BlockSparseMatrix
 - BlockSparseMatrix::ToCompressedRowSparseMatrix is performed in a
   direct sort-less way

Change-Id: Ib951fda1b9394050e2c47a9721172c5e3c674801
2023-04-18 01:34:30 +03:00
Sameer Agarwal d340f81bd0 Clang Tidy fixes
Change-Id: I51429acbf2a7b81605a2fd03a6b7c10317984674
2023-04-10 16:43:30 -07:00
Dmitriy Korchemkin 54ad3dd03c Reorganize ParallelFor source files
Change-Id: Ic4941919e59210b48e447cbb61e539200c8c89df
2023-04-11 00:33:55 +03:00
Sameer Agarwal ba360ab074 Change the value of BlockRandomAccessSparseMatrix::kMaxRowBlocks
1. Rename it to kRowShift.
2. Make it a static constexpr.
3. Change it to 2^32, which should allow for easier bit arithmetic
   for the compiler than the previously used value of 10000000.
4. Change the name of the two associated private methods from
   IntPairToLong to IntPairToInt64 and LongToIntPair to Int64ToIntPair.

Change-Id: I54d61bcf1121079b222ef518324de5cffc1be064
2023-04-09 14:04:58 -07:00
Sergiu Deitsch 0315c6ca9a Provide DynamicAutoDiffCostFunction deduction guide
The deduction guide allows to avoid repeating the CostFunctor type.

Change-Id: I2285de37071006a97f89988baa9b7054d82dae86
2023-03-08 23:46:33 +01:00
MaximSmolskiy 3cdfae110f Replace Hertzberg mentions with citations
Change-Id: Ife1b9c2b321593912bc18a5c57c39fbda27da51f
2023-03-04 04:10:45 +03:00
MaximSmolskiy 0af38a9fc2 Fix typos in documentation
Change-Id: I748c9c3d6cd1c5906daa9f68aea92d8ad536bf3b
2023-03-04 03:56:54 +03:00
MaximSmolskiy 4c969a6c1c Improve image of loss functions shape
Change-Id: Id226323a736038fb3c9b9009448e5ece12dc63ee
2023-03-03 04:07:01 +03:00
MaximSmolskiy b54f05b8ee Add missing TukeyLoss to documentation
Change-Id: I0b0a84c0a21b672f0414eb6b2283bb27e06cd266
2023-03-03 00:19:51 +03:00
Dmitriy Korchemkin 8bf4a2f42c Inexact check for ParallelAssign test
Change-Id: I3880f51868c55bce2901e7414cf5109385cccee7
2023-01-31 17:13:19 +03:00
Alexander Ivanov 9cddce73a5 Explicit conversions from long to int in benchmarks (for num_threads)
Change-Id: I175328a890efe79c97180be03997324532d4c7a7
2023-01-29 01:56:23 +00:00