Commit Graph

1298 Commits

Author SHA1 Message Date
Sameer Agarwal 5a30cae583 Preparing for 2.2.0rc1
1. Add a version history
2. Update copyright years across the code base
3. Run format_all.sh
4. Update version strings from 2.1.0 to 2.2.0 in the docs and
   elsewhere.

Change-Id: I46d8d479d54bd6002d532785e67342106e73c9ac
2023-09-21 11:23:38 -07:00
Mark Shachkov 9cca671273 Enable compatibility with SuiteSparse 7.2.0
Change-Id: I072dc3f7c245fc2ebbdffed715ac4def20f7dccd
2023-09-17 20:57:43 +02:00
Sergiu Deitsch a57e35bbab Require at least CMake 3.16
Given we no longer support Ubuntu 18.04 due to packaged GCC lacking
C++17 support we can bump the minimum required CMake version to the one
provided by Ubuntu 20.04 which is CMake 3.16. Consequently, this allows
to drop some of the legacy CMake logic.

Change-Id: I1f05d4c5681d10aa7faa0800ef4a803be2f5b7dd
2023-09-12 19:33:00 +02:00
Sergiu Deitsch 863db948f3 Eliminate macOS sprintf warning
AppleClang 14.0.0.14000029 warns about a potential security problem
while invoking the sprintf C function:

    internal/ceres/fixed_array_test.cc:469:3: warning: 'sprintf' is deprecated: This function is provided for compatibility reasons only.  Due to security concerns inherent in the design of sprintf(3), it is highly recommended that you use snprintf(3) instead. [-Wdeprecated-declarations]
      sprintf(buf.data(), "foo");  // NOLINT(runtime/printf)
      ^
    /Applications/Xcode_14.2.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX13.1.sdk/usr/include/stdio.h:188:1: note: 'sprintf' has been explicitly marked deprecated here
    __deprecated_msg("This function is provided for compatibility reasons only.  Due to security concerns inherent in the design of sprintf(3), it is highly recommended that you use snprintf(3) instead.")
    ^
    /Applications/Xcode_14.2.app/Contents/Developer/Platforms/MacOSX.platform/Developer/SDKs/MacOSX13.1.sdk/usr/include/sys/cdefs.h:215:48: note: expanded from macro '__deprecated_msg'
            #define __deprecated_msg(_msg) __attribute__((__deprecated__(_msg)))

Replace sprintf by snprintf to avoid this deprecation warning.

Change-Id: I6870c0bd4e390388d1d7bcec082cee272b234eba
2023-09-11 22:51:28 +02:00
Dmitriy Korchemkin 5e4b22f7fc Update CudaSparseMatrix class
- Perform temporary buffer size estimation only once
- Allow construction from existing buffers with col/row structure

Change-Id: I73c291328f1e8ed9184aba5d7058df71cbc6a15d
2023-08-31 18:56:44 +00:00
Sameer Agarwal de62bf2204 Two minor fixes
1. In cuda_sparse_matrix.cc fix the order of fields in the initializer list.
2. Move a line of code to the ifdef branch which will use it.

Change-Id: If32ea14a287f845c1740e6f726c1007e86a4eeca
2023-08-20 17:53:10 +00:00
Dmitriy Korchemkin ec4907399a Fix block-sparse to crs conversion on windows
Change-Id: I2aabdb68afc8152c5d8da157baf22faf8d8f80cf
2023-08-16 16:16:41 +00:00
Dmitriy Korchemkin 799ee91bbf Fix check in CompressedRowJacobianWriter::CreateJacobian()
Change-Id: I8c8d418f14a9e489a94b5117460e739f65e36a96
2023-08-16 11:14:42 +00:00
Sameer Agarwal ee90f6cbf7 Detect large Jacobians and return failure instead of crashing.
Detect when the number of non-zeros overflows when constructing
BlockSparseMatrix and CompressedRowSparseMatrix and return
with an error message instead of crashing.

Change-Id: I45e102f7c0519eef441ce0586b7adf96e4a954a9
2023-08-11 17:16:28 -07:00
Sameer Agarwal 310a252fb6 Deal with infinite initial cost correctly.
Previously it could be the case that a residual block could return
a residual whose squared norm overflows and generates an infinity
which we did not detect. This would then lead to the trust region
minimizer incorrectly terminating indicating convergence while
generating a cost delta of NaN.

This change adds a check for that and also does two minor cosmetic
changes.

1. Reduce the level of nesting in program_evaluator.h by adding
   an early return.
2. The error message when IterationZero fails now says that the
   Initial residual and Jacobian failed, to indicate that the
   optimizer had no chance to do any work.

Fixes https://github.com/ceres-solver/ceres-solver/issues/988

Thanks to @Ashray-g for reporting this.

Change-Id: I52ae7627a66f637135209dbb2e42935b52c8bc77
2023-08-04 18:35:25 -07:00
Dmitriy Korchemkin 75bacedf7d CUDA partitioned matrix view
Converts BlockSparseMatrix into two instances of CudaSparseMatrix,
corresponding to left and right sub-matrix.

Values of submatrix E are always just copied as-is, and values of
submatrix F are copied if each row-block of F submatrix satisfies
at least one of the following conditions:
 - There is atmost one cell in row-block
 - Row block has height of 1 row
Otherwise, indices of values in CRS order corresponding to value indices
in block-sparse order are computed on-the-fly.

Change-Id: I14eee00c36ee74b6b83fc85927907641383abfc7
2023-07-14 18:12:20 +00:00
Sergiu Deitsch fd6197ce0e Fixed GCC 13.1 compilation errors
Change-Id: I4c0ad38c899a9e5988e8e9f7cb870116aa8233eb
2023-06-10 17:37:51 +02:00
Dmitriy Korchemkin df97a8f059 Improve support of older CUDA toolkit versions
- Provides fall-back for older versions of CUDA toolkit
 - Using older versions of CUDA toolkit might result in
   over-synchronization

Change-Id: I545e6625d2342be30cb759b90bda379e555d7370
2023-05-31 22:45:35 +03:00
Dmitriy Korchemkin 085214ea74 Fix test CompressedRowSparseMatrix.Transpose
Change-Id: I653b4619d518f767e57f04136117d7fe47f33ca1
2023-05-30 16:03:57 +03:00
Dmitriy Korchemkin d880df09f9 Match new[] with delete[] in BSM
Change-Id: If78911c9570ce6a9039192501e6da7db3974293a
2023-05-30 11:28:33 +03:00
Dmitriy Korchemkin bdee4d6172 Block-sparse to CRS conversion using block-structure
Instead of pre-computing pemutation from block-sparse to CRS order,
index of value in CRS matrix is computed in the process of updating
values using block-sparse structure.

When it is possible to update values via a simple host-to-device copy,
block-sparse structure on GPU is discarded after computing CRS
structure.

Computing index is significantly slower than using pre-computed
permutation, but is still hidden by host-to-device transfer.

On problems from BAL dataset this results into reduction of extra
gpu memory consumption from 33% (permutation stored as 32-bit indices)
to ~10% for storing block-sparse structure.

Benchmark results:

======================= CUDA Device Properties ======================
Cuda version         : 11.8
Device ID            : 0
Device name          : NVIDIA GeForce RTX 2080 Ti
Total GPU memory     :  11012 MiB
GPU memory available :  10852 MiB
Compute capability   : 7.5
Warp size            : 32
Max threads per block: 1024
Max threads per dim  : 1024 1024 64
Max grid size        : 2147483647 65535 65535
Multiprocessor count : 68
====================================================================
Running ./bin/evaluation_benchmark
Run on (112 X 3200 MHz CPU s)
CPU Caches:
  L1 Data 32 KiB (x56)
  L1 Instruction 32 KiB (x56)
  L2 Unified 1024 KiB (x56)
  L3 Unified 39424 KiB (x2)
Load Average: 24.58, 11.75, 8.52

-----------------------------------------------------------------------
Benchmark                                                          Time
-----------------------------------------------------------------------
Using on-the-fly computation of CRS index corresponding to block-sparse
index:

JacobianToCRS<g/final/problem-4585-1324582-pre.txt>             1607 ms
JacobianToCRSView<g/final/problem-4585-1324582-pre.txt>          564 ms
JacobianToCRSMatrix<g/final/problem-4585-1324582-pre.txt>       2226 ms
JacobianToCRSViewUpdate<g/final/problem-4585-1324582-pre.txt>    228 ms
JacobianToCRSMatrixUpdate<g/final/problem-4585-1324582-pre.txt>  400 ms

Using precomputed permutation:
JacobianToCRS</final/problem-4585-1324582-pre.txt>              1656 ms
JacobianToCRSView</final/problem-4585-1324582-pre.txt>           553 ms
JacobianToCRSMatrix</final/problem-4585-1324582-pre.txt>        2255 ms
JacobianToCRSViewUpdate</final/problem-4585-1324582-pre.txt>     228 ms
JacobianToCRSMatrixUpdate</final/problem-4585-1324582-pre.txt>   406 ms

Performance of JacobianToCRSViewUpdate is still limited by
host-to-device transfer, and JacobianToCRSView is faster than computing
CRS structure on CPU.

Change-Id: Ifb6910fb01ae6071400d36c277846fadc5857964
2023-05-26 01:12:47 +03:00
Sameer Agarwal 0f9de3daf4 Use page locked memory in BlockSparseMatrix
If using CUDA_SPARSE for an iterative solve on the GPU,
allocate the values array in BlockSparseMatrix to make copying
to the GPU faster.

Change-Id: I63c1d2512babd74fc275b277ac8c3eabf3ec1144
2023-05-15 12:25:38 -07:00
Dmitriy Korchemkin e7bd72d41e Permutation-based conversion from block-sparse to crs
Change-Id: Ic33a6476c033187dff61886deb6d1761524943f0
2023-05-12 03:33:25 +03:00
Dmitriy Korchemkin abbc4e7974 Explicitly compute number of non-zeros in row
Change-Id: Iaebaf6d23d33dbbbe5a7c7240319bf28cf2bdd3a
2023-05-04 21:16:39 +03:00
Hs293Go 96fdfd2e7a Implement tests for Euler conversion with jets
https://github.com/ceres-solver/ceres-solver/issues/965

Change-Id: I6bdabf5ee09c000e49ea1757fde724d4bc4ccb92
2023-04-27 16:42:25 -04:00
Sameer Agarwal a5e745d4e1 ClangTidy fixes
Change-Id: I43badfd6706a44148989ac5621134453bad28133
2023-04-18 06:37:52 -07:00
Dmitriy Korchemkin 77ad8bb4e5 Change storage in BlockRandomAccessSparseMatrix
- TripletSparseMatrix in BlockRandomAccessSparseMatrix is replaced with
   BlockSparseMatrix
 - BlockSparseMatrix::ToCompressedRowSparseMatrix is performed in a
   direct sort-less way

Change-Id: Ib951fda1b9394050e2c47a9721172c5e3c674801
2023-04-18 01:34:30 +03:00
Sameer Agarwal d340f81bd0 Clang Tidy fixes
Change-Id: I51429acbf2a7b81605a2fd03a6b7c10317984674
2023-04-10 16:43:30 -07:00
Dmitriy Korchemkin 54ad3dd03c Reorganize ParallelFor source files
Change-Id: Ic4941919e59210b48e447cbb61e539200c8c89df
2023-04-11 00:33:55 +03:00
Sameer Agarwal ba360ab074 Change the value of BlockRandomAccessSparseMatrix::kMaxRowBlocks
1. Rename it to kRowShift.
2. Make it a static constexpr.
3. Change it to 2^32, which should allow for easier bit arithmetic
   for the compiler than the previously used value of 10000000.
4. Change the name of the two associated private methods from
   IntPairToLong to IntPairToInt64 and LongToIntPair to Int64ToIntPair.

Change-Id: I54d61bcf1121079b222ef518324de5cffc1be064
2023-04-09 14:04:58 -07:00
Sergiu Deitsch 0315c6ca9a Provide DynamicAutoDiffCostFunction deduction guide
The deduction guide allows to avoid repeating the CostFunctor type.

Change-Id: I2285de37071006a97f89988baa9b7054d82dae86
2023-03-08 23:46:33 +01:00
Dmitriy Korchemkin 8bf4a2f42c Inexact check for ParallelAssign test
Change-Id: I3880f51868c55bce2901e7414cf5109385cccee7
2023-01-31 17:13:19 +03:00
Alexander Ivanov 9cddce73a5 Explicit conversions from long to int in benchmarks (for num_threads)
Change-Id: I175328a890efe79c97180be03997324532d4c7a7
2023-01-29 01:56:23 +00:00
Alexander Ivanov 5e787ab700 Using const ints in jacobian writers
Change-Id: I90ebb1956a15eddad2758bed7acaf3a2b90ce7c3
2023-01-28 03:42:03 +00:00
Sameer Agarwal e269b64f55 More ClangTidy fixes
Change-Id: I8288a61cf0db80efd7fd4862a555cc717797e8ba
2023-01-17 13:37:13 -08:00
Alexander Ivanov f4eb768e0b Using int64 in file.cc. Fixing compilation error in array_utils.cc
Change-Id: Ie38e257df59bf1884ef431d5d25b3c5ac0dde7a1
2023-01-17 18:59:36 +00:00
Alexander Ivanov 74a0f0d246 Using int64_t for sizes and indexes in array utils
Change-Id: I61177d868850c4538adbd77e413238d9cb2d9248
2023-01-17 15:27:56 +00:00
Sameer Agarwal 749a442d97 Clang-Tidy fixes
Change-Id: I58900a452591315a39754b329e94b315c34926cd
2023-01-16 07:38:05 -08:00
Sameer Agarwal c4ba975aea Fix rotation_test.cc to work with older versions of Eigen
Change-Id: I699e4b45f9a1cc09e72f8140e67c800c3bcef8e0
2023-01-14 06:08:10 -08:00
Sameer Agarwal 9602ed7b76 ClangFormat changes
Change-Id: I88c9e38b0450aed26c60e1dd54964ab6571e3eef
2023-01-14 05:54:24 -08:00
Sameer Agarwal 79a554ffcf Fix a bug in QuaternionRotatePoint.
In https://ceres-solver-review.git.corp.google.com/c/ceres-solver/+/23802

the computation of the norm of a quaternion

scale = 1/sqrt(q[0] * q[0] + q[1] * q[1] + q[2] * q[2] + q[3] * q[3]);

was replaced by

scale = 1/hypot(q[0], q[1], hypot(q[2], q[3]));

while this appear to be a more accurate computation because of the
use of hypot which can handle over and underflow it introduces a
bug for the case where q[2] = q[3] = 0.

While the hypot(q[2], q[3]) == 0 as scalars, if q[2] and q[3] are
jets, then the derivative will be NaN. Which means that even though
q[0] or q[1] is non-zero and the norm of the quaternion is non-zero,
and the resulting derivative is finite, this way of computing the
scale will produce nans in the derivative of scale.

The following quaternion will replicate the problem described above.

using Jet = ceres::Jet<double, 4>;
std::array<Jet, 4> quaternion = {Jet(1.0, 0), Jet(0.0, 1), Jet(0.0, 2), Jet(0.0, 3)};

This CL reverts the change to QuaternionRotatePoint and
adds a test for it.

Thanks to Jonathan Taylor for reproducing this bug.

Change-Id: I0fbbcc77d6945a38563d82efba4429f4b5278cd5
2023-01-13 11:37:27 -08:00
Alexander Ivanov f1113c08ab Commenting unused parameters for better readibility
Change-Id: Idc285fa68ba787636a69a3ea3350e0282b9f8569
2023-01-12 17:25:42 +00:00
Alexander Ivanov 772d927e19 Replacing old style typedefs with new style usings
Change-Id: I85d353708fc431df8312a2337d0508df6aee071f
2023-01-12 09:09:35 +00:00
Alexander Ivanov 53df5ddcfd Removing using std::...
Change-Id: I584402e2a34869183c1d59071a15d97b216c52fb
2023-01-11 16:51:38 +00:00
Sergiu Deitsch cb6b306623 Use hypot to compute the L^2 norm
Change-Id: I908eaaa279452aa16346dfd3f25aac53685e3172
2023-01-05 21:44:53 +01:00
Sameer Agarwal 73d95b03fa Clang-Tidy fixes
Change-Id: Ib391357fe365e2321e75b106ffac5f832abbb8fa
2023-01-03 11:59:43 -08:00
Joydeep Biswas 51d52c3ea5 Correct epsilon in CUDA QR test changed by last commit.
Change-Id: I84d40b1df36a972eb4b17e6cdfc803a5367cb48a
2022-12-28 12:37:51 -06:00
Joydeep Biswas 546f5337bd Updates to CUDA dense linear algebra tests
* Use relative error instead of absolute error.
* Update tolerance to account for embedded GPUs such as the Jetson TX2.

Change-Id: I05742ecbfd915e797fcbc4137acc5d1b513cd465
2022-12-28 12:20:20 -06:00
Sameer Agarwal 19ab2c1793 BlockRandomAccessMatrix Refactor
1. Add threading to all three subclasses of BlockRandomAccessMatrix.
   i.e. BlockRandomAccessDenseMatrix, BlockRandomAccessSparseMatrix
   and BlockRandomAccessDenseMatrix.

   For BlockRandomAccessDenseMatrix and BlockRandomAccessSparseMatrix
   this just means SetZero is parallelized. Which by itself is no
   big deal, but by doing so, the constructor for all three subclasses
   become uniform.

   BlockRandomAccessSparseMatrix::SymmetricRightMultiplyAndAccumulate
   maybe threaded in the future if needed.

   BlockRandomAccessDiagonalMatrix is the biggest beneficiary. SetZero
   Invert and RightMultiplyAndAccumulate are all threaded now.

2. Change the storage in BlockRandomAccessDiagonalMatrix from
   TripletSparseMatrix to CompressedRowSparseMatrix. This has no
   performance implications since we do not really use the capabilities
   of the underlying matrix indexing representation. This is a forward
   looking change when we decide to transfer this matrix to the GPU,
   a CompressedRowSparseMatrix will save on a data conversion.

3. Use std::unique_ptr as needed and eliminate the need for custom
   destructors.

4. Modify CompressedRowSparseMatrix::CreateBlockDiagonalMatrix to
   take a nullptr as the data vector.

Fixes https://github.com/ceres-solver/ceres-solver/issues/936
Fixes https://github.com/ceres-solver/ceres-solver/issues/935

Change-Id: Ia6487f2d924fbe669835bdcc38abf2b451bda4ee
2022-12-23 06:45:28 -08:00
Sameer Agarwal a3a062d72c Add a missing header
Change-Id: I26705eb557a8e21b6acc7d18a01471d38a3c1c23
2022-12-19 06:42:35 -08:00
Sameer Agarwal 2b88bedb28 Remove unused variables
Change-Id: Ieaecce32d8a94c61876e097b957627d689e72fe7
2022-12-19 06:27:59 -08:00
Sameer Agarwal 9beea728f6 Fix a bug in CoordinateDescentMinimizer
CoordinateDescentMinimizer optimizes one parameter block at a time.
To do this, it manipulates the parameter block object. It was doing
so inconsistently, where the tangent space offset was being set to
zero but the ambient state offset was not being set to zero. This
did not cause problems because these offsets were not really being
used inside the CoordinateDescentMinimizer. However the recent
change which parallelizes Program::Plus uncovered this bug.

The reason this bug was not caught was because, CoordinateDescentMinimizer
does not have any tests. I will fix this shortly, but in the interim
to unbreak inner iterations at head, this small change should go in.

Change-Id: I55d2698e8509f9cb5751e7a5180427129d86e720
2022-12-17 17:32:03 -08:00
Sameer Agarwal 8e5d83f07d ClangFormat and ClangTidy changes
Change-Id: Ib457dcc55ffb405aeaeac711c20bd9217b32f90e
2022-12-17 17:29:42 -08:00
Dmitriy Korchemkin b158515089 Parallel operations on vectors
Main focus of this change is to parallelize remaining operations (most of them
are operations on vectors) in code-path utilized with iterative Schur
complement.

Parallelization is handled using lazy evaluation of Eigen expressions.

On linux pc with intel 8176 processor parallelization of vector operations has
the following effect:

Running ./bin/parallel_vector_operations_benchmark
Run on (112 X 3200.32 MHz CPU s)
CPU Caches:
  L1 Data 32 KiB (x56)
  L1 Instruction 32 KiB (x56)
  L2 Unified 1024 KiB (x56)
  L3 Unified 39424 KiB (x2)
Load Average: 3.30, 8.41, 11.82
-----------------------------------
Benchmark                      Time
-----------------------------------
SetZero                 10009532 ns
SetZeroParallel/1       10024139 ns
...
SetZeroParallel/16        877606 ns

Negate                   4978856 ns
NegateParallel/1         5145413 ns
...
NegateParallel/16         721823 ns

Assign                  10731408 ns
AssignParallel/1        10749944 ns
...
AssignParallel/16        1829381 ns

D2X                     15214399 ns
D2XParallel/1           15623245 ns
...
D2XParallel/16           2687060 ns

DivideSqrt               8220050 ns
DivideSqrtParallel/1     9088467 ns
...
DivideSqrtParallel/16     905569 ns

Clamp                    3502010 ns
ClampParallel/1          4507897 ns
...
ClampParallel/16          759576 ns

Norm                     4426782 ns
NormParallel/1           4442805 ns
...
NormParallel/16           430290 ns

Dot                      9023276 ns
DotParallel/1            9031304 ns
...
DotParallel/16           1157267 ns

Axpby                   14608289 ns
AxpbyParallel/1         14570825 ns
...
AxpbyParallel/16         2672220 ns
-----------------------------------

Multi-threading of vector operations in ISC and program evaluation results into
the following improvement:

Running ./bin/evaluation_benchmark
--------------------------------------------------------------------------------------
Benchmark                                                               this   2fd81de
--------------------------------------------------------------------------------------
Residuals<problem-13682-4456117-pre.txt>/1                           4136 ms   4292 ms
Residuals<problem-13682-4456117-pre.txt>/2                           2919 ms   2670 ms
Residuals<problem-13682-4456117-pre.txt>/4                           2065 ms   2198 ms
Residuals<problem-13682-4456117-pre.txt>/8                           1458 ms   1609 ms
Residuals<problem-13682-4456117-pre.txt>/16                          1152 ms   1227 ms

ResidualsAndJacobian<problem-13682-4456117-pre.txt>/1               19759 ms  20084 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/2               10921 ms  10977 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/4                6220 ms   6941 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/8                3490 ms   4398 ms
ResidualsAndJacobian<problem-13682-4456117-pre.txt>/16               2277 ms   3172 ms

Plus<problem-13682-4456117-pre.txt>/1                                 339 ms    322 ms
Plus<problem-13682-4456117-pre.txt>/2                                 220 ms
Plus<problem-13682-4456117-pre.txt>/4                                 128 ms
Plus<problem-13682-4456117-pre.txt>/8                                78.0 ms
Plus<problem-13682-4456117-pre.txt>/16                               49.8 ms

ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/1       2434 ms   2478 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/2       2706 ms   2688 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/4       1430 ms   1548 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/8        742 ms    883 ms
ISCRightMultiplyAndAccumulate<problem-13682-4456117-pre.txt>/16       438 ms    555 ms

ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/1   2438 ms   2481 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/2   2565 ms   2790 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/4   1434 ms   1551 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/8    765 ms    892 ms
ISCRightMultiplyAndAccumulateDiag<problem-13682-4456117-pre.txt>/16   435 ms    559 ms

JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/1           1278 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/2           1555 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/4            833 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/8            459 ms
JacobianSquaredColumnNorm<problem-13682-4456117-pre.txt>/16           250 ms

JacobianScaleColumns<problem-13682-4456117-pre.txt>/1                1468 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/2                1871 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/4                 957 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/8                 528 ms
JacobianScaleColumns<problem-13682-4456117-pre.txt>/16                294 ms

End-to-end improvements with bundle_adjuster invoked with
./bin/bundle_adjuster --num_threads 28 --num_iterations 40 \
                      --linear_solver iterative_schur \
                      --preconditioner jacobi --input
---------------------------------------------
Problem                         this  2fd81de
---------------------------------------------
problem-13682-4456117-pre.txt  508.6    892.7
problem-1778-993923-pre.txt    763.8   1129.9
problem-1723-156502-pre.txt      6.3     14.4
problem-356-226730-pre.txt      76.3    116.2
problem-257-65132-pre.txt       38.6     52.0

Change-Id: Ie31cc5015f13fa479c16ffb5ce48c9b880990d49
2022-12-17 02:52:27 +03:00
Dmitriy Korchemkin 2fd81de12d Add build configuration with CUDA on Linux
Change-Id: I3144a44692a7a129857b65ed84fb2a5637b25b5d
2022-11-29 18:51:13 +03:00