Sean McBride
|
d70b4864d9
|
issue #2581: review and cleanup of compiler version checks
|
2023-01-17 18:58:34 +00:00 |
|
Antonio Sánchez
|
8588d8c74b
|
Correct pnegate for floating-point zero.
|
2022-11-15 18:07:23 +00:00 |
|
Charles Schlosser
|
82b152dbe7
|
Add signbit function
|
2022-11-04 00:31:20 +00:00 |
|
Antonio Sánchez
|
e5794873cb
|
Replace assert with eigen_assert.
|
2022-10-04 17:11:23 +00:00 |
|
Rasmus Munk Larsen
|
c475228b28
|
Vectorize atan() for double.
|
2022-10-01 01:49:30 +00:00 |
|
Rasmus Munk Larsen
|
13b69fc1b0
|
Try to reduce compilation time/memory for GEBP kernel using EIGEN_IF_CONSTEXPR
|
2022-09-23 20:09:42 +00:00 |
|
Rasmus Munk Larsen
|
7b2901e2aa
|
Add vectorized integer division for int32 with AVX512, AVX or SSE.
|
2022-09-21 00:27:23 +00:00 |
|
Rasmus Munk Larsen
|
bd393e15c3
|
Vectorize acos, asin, and atan for float.
|
2022-08-29 19:49:33 +00:00 |
|
Charles Schlosser
|
e5af9f87f2
|
Vectorize pow for integer base / exponent types
|
2022-08-29 19:23:54 +00:00 |
|
Matthew Sterrett
|
7a3b667c43
|
Add support for AVX512-FP16 for vectorizing half precision math
|
2022-08-17 18:15:21 +00:00 |
|
Matthew Sterrett
|
39fcc89798
|
Removed unnecessary checks for FP16C
|
2022-08-16 18:14:41 +00:00 |
|
Antonio Sánchez
|
2cf4d18c9c
|
Disable AVX512 GEMM kernels by default.
|
2022-07-20 21:22:48 +00:00 |
|
b-shi
|
4a56359406
|
Add option to disable avx512 GEBP kernels
|
2022-07-18 17:59:09 +00:00 |
|
b-shi
|
37673ca1bc
|
AVX512 TRSM kernels use alloca if EIGEN_NO_MALLOC requested
|
2022-06-17 18:05:26 +00:00 |
|
Shi, Brian
|
28812d2ebb
|
AVX512 TRSM Kernels respect EIGEN_NO_MALLOC
|
2022-06-07 11:28:42 -07:00 |
|
aaraujom
|
8fbb76a043
|
Fix build issues with MSVC for AVX512
|
2022-06-03 14:55:40 +00:00 |
|
aaraujom
|
d49ede4dc4
|
Add AVX512 s/dgemm optimizations for compute kernel (2nd try)
|
2022-05-28 02:00:21 +00:00 |
|
Antonio Sánchez
|
9b9496ad98
|
Revert "Add AVX512 optimizations for matrix multiply"
This reverts commit 25db0b4a82
|
2022-05-13 18:50:33 +00:00 |
|
aaraujom
|
25db0b4a82
|
Add AVX512 optimizations for matrix multiply
|
2022-05-12 23:41:19 +00:00 |
|
Antonio Sánchez
|
07db964bde
|
Restrict new AVX512 trsm to AVX512VL, rename files for consistency.
|
2022-04-14 16:58:32 +00:00 |
|
b-shi
|
0611f7fff0
|
Add missing explicit reinterprets
|
2022-03-23 21:10:26 +00:00 |
|
Antonio Sánchez
|
4451823fb4
|
Fix ODR violation in trsm.
|
2022-03-20 15:56:53 +00:00 |
|
Antonio Sánchez
|
9a14d91a99
|
Fix AVX512 builds with MSVC.
|
2022-03-18 16:04:53 +00:00 |
|
b-shi
|
518fc321cb
|
AVX512 Optimizations for Triangular Solve
|
2022-03-16 18:04:50 +00:00 |
|
Sean McBride
|
f1b9692d63
|
Removed EIGEN_UNUSED decorations from many functions that are in fact used
|
2022-03-03 20:19:33 +00:00 |
|
Antonio Sánchez
|
9c07e201ff
|
Modified sqrt/rsqrt for denormal handling.
|
2022-03-02 17:20:47 +00:00 |
|
Antonio Sánchez
|
19c39bea29
|
Fix mixingtypes for g++-11.
|
2022-02-25 19:28:10 +00:00 |
|
Rasmus Munk Larsen
|
979fdd58a4
|
Add generic fast psqrt and prsqrt impls and make them correct for 0, +Inf, NaN, and negative arguments.
|
2022-02-05 00:20:13 +00:00 |
|
Antonio Sánchez
|
e7f4a901ee
|
Define EIGEN_HAS_AVX512_MATH in PacketMath.
|
2022-02-04 22:25:52 +00:00 |
|
Antonio Sánchez
|
96da541cba
|
Fix AVX512 math function consistency, enable for ICC.
|
2022-02-04 19:35:18 +00:00 |
|
Rasmus Munk Larsen
|
51311ec651
|
Remove inline assembly for FMA (AVX) and add remaining extensions as packet ops: pmsub, pnmadd, and pnmsub.
|
2022-01-26 04:25:41 +00:00 |
|
Rasmus Munk Larsen
|
ea2c02060c
|
Add reciprocal packet op and fast specializations for float with SSE, AVX, and AVX512.
|
2022-01-21 23:49:18 +00:00 |
|
Ilya Tokar
|
a0fc640c18
|
Add support for packets of int64 on x86
|
2022-01-21 19:55:23 +00:00 |
|
Kolja Brix
|
8d81a2339c
|
Reduce usage of reserved names
|
2022-01-10 20:53:29 +00:00 |
|
Kolja Brix
|
afa616bc9e
|
Fix some typos found
|
2021-09-23 15:22:00 +00:00 |
|
Antonio Sanchez
|
3c724c44cf
|
Fix strict aliasing bug causing product_small failure.
Packet loading is skipped due to aliasing violation, leading to nullopt matrix
multiplication.
Fixes #2327.
|
2021-09-17 21:09:34 +00:00 |
|
Rasmus Munk Larsen
|
7b975acb1f
|
Remove unused variable.
|
2021-09-16 20:27:13 +00:00 |
|
Rasmus Munk Larsen
|
92849d814b
|
Remove unused variable.
|
2021-09-16 20:21:31 +00:00 |
|
Rasmus Munk Larsen
|
d7d0bf832d
|
Issue an error in case of direct inclusion of internal headers.
|
2021-09-10 19:12:26 +00:00 |
|
Antonio Sanchez
|
3d4ba855e0
|
Fix AVX integer packet issues.
Most are instances of AVX2 functions not protected by
`EIGEN_VECTORIZE_AVX2`. There was also a missing semi-colon
for AVX512.
|
2021-09-01 14:14:43 -07:00 |
|
Jakub Lichman
|
dc5b1f7d75
|
AVX512 and AVX2 support for Packet16i and Packet8i added
|
2021-08-25 19:38:23 +00:00 |
|
Gauri Deshpande
|
e6a5a594a7
|
remove denormal flushing in fp32tobf16 for avx & avx512
|
2021-08-09 22:15:21 +00:00 |
|
Rasmus Munk Larsen
|
9312a5bf5c
|
Implement a generic vectorized version of Smith's algorithms for complex division.
|
2021-07-01 23:31:12 +00:00 |
|
Rasmus Munk Larsen
|
52a5f98212
|
Get rid of code duplication for conj_helper. For packets where LhsType=RhsType a single generic implementation suffices. For scalars, the generic implementation of pconj automatically forwards to numext::conj, so much of the existing specialization can be avoided. For mixed types we still need specializations.
|
2021-06-24 15:47:48 -07:00 |
|
Jakub Lichman
|
d87648a6be
|
Tests added and AVX512 bug fixed for pcmp_lt_or_nan
|
2021-04-25 20:58:56 +00:00 |
|
Jakub Lichman
|
2b1dfd1ba0
|
HasExp added for AVX512 Packet8d
|
2021-04-20 19:07:58 +00:00 |
|
Antonio Sanchez
|
1d79c68ba0
|
Fix ldexp for AVX512 (#2215)
Wrong shuffle was used. Need to interleave low/high halves with a
`permute` instruction.
Fixes #2215.
|
2021-04-20 16:25:22 +00:00 |
|
Christoph Hertzberg
|
69a4f70956
|
Revert "Uses _mm512_abs_pd for Packet8d pabs"
This reverts commit f019b97aca
|
2021-03-23 18:52:19 +00:00 |
|
Steve Bronder
|
f019b97aca
|
Uses _mm512_abs_pd for Packet8d pabs
|
2021-03-18 15:47:52 +00:00 |
|
Antonio Sanchez
|
7ff0b7a980
|
Updated pfrexp implementation.
The original implementation fails for 0, denormals, inf, and NaN.
See #2150
|
2021-02-17 02:23:24 +00:00 |
|