Jakub Lichman
dc5b1f7d75
AVX512 and AVX2 support for Packet16i and Packet8i added
2021-08-25 19:38:23 +00:00
Han-Kuan Chen
ab28419298
optimize predux if architecture is aarch64
2021-08-25 19:18:54 +00:00
Antonio Sanchez
2cc6ee0d2e
Add missing PPC packet comparisons.
...
This is to fix the packetmath tests on the ppc pipeline.
2021-08-17 07:42:04 -07:00
Chip-Kerchner
8dcf3e38ba
Fix unaligned loads in ploadLhs & ploadRhs for P8.
2021-08-16 20:28:22 -05:00
Chip-Kerchner
e07227c411
Reverse compare logic in F32ToBf16 since vec_cmpne is not available in Power8 - now compiles for clang10 default (P8).
2021-08-13 11:21:28 -05:00
Chip Kerchner
66499f0f17
Get rid of used uninitialized warnings for EIGEN_UNUSED_VARIABLE in gcc11+
2021-08-12 21:38:54 +00:00
ChipKerchner
413bc491f1
Fix errors on older compilers (gcc 7.5 - lack of vec_neg, clang10 - can not use const pointers with vec_xl).
2021-08-10 15:03:18 -05:00
Gauri Deshpande
e6a5a594a7
remove denormal flushing in fp32tobf16 for avx & avx512
2021-08-09 22:15:21 +00:00
derekjchow
66ca41bd47
Add support for vectorizing logical comparisons.
2021-07-23 20:07:48 +00:00
Rasmus Munk Larsen
7b35638ddb
Fix breakage of conj_helper in conjunction with custom types introduced in !537 .
2021-07-02 20:42:15 +00:00
Rasmus Munk Larsen
bbfc4d54cd
Use padd instead of +.
2021-07-02 02:51:48 +00:00
Rasmus Munk Larsen
9312a5bf5c
Implement a generic vectorized version of Smith's algorithms for complex division.
2021-07-01 23:31:12 +00:00
Chip Kerchner
91e99ec1e0
Create the ability to disable the specialized gemm_pack_rhs in Eigen (only PPC) for TensorFlow
2021-06-30 23:05:04 +00:00
大河メタル
c81da59a25
Correct declarations for aarch64-pc-windows-msvc
2021-06-30 04:09:46 +00:00
Rasmus Munk Larsen
5aebbe9098
Get rid of redundant pabs instruction in complex square root.
2021-06-29 23:26:15 +00:00
Rohit Santhanam
2d132d1736
Commit 52a5f982 broke conjhelper functionality for HIP GPUs.
...
This commit addresses this.
2021-06-25 19:28:00 +00:00
Rasmus Munk Larsen
bffd267d17
Small cleanup: Get rid of the macros EIGEN_HAS_SINGLE_INSTRUCTION_CJMADD and CJMADD, which were effectively unused, apart from on x86, where the change results in identically performing code.
2021-06-24 18:52:17 -07:00
Rasmus Munk Larsen
52a5f98212
Get rid of code duplication for conj_helper. For packets where LhsType=RhsType a single generic implementation suffices. For scalars, the generic implementation of pconj automatically forwards to numext::conj, so much of the existing specialization can be avoided. For mixed types we still need specializations.
2021-06-24 15:47:48 -07:00
Antonio Sanchez
12e8d57108
Remove pset, replace with ploadu.
...
We can't make guarantees on alignment for existing calls to `pset`,
so we should default to loading unaligned. But in that case, we should
just use `ploadu` directly. For loading constants, this load should hopefully
get optimized away.
This is causing segfaults in Google Maps.
2021-06-16 18:41:17 -07:00
Chip-Kerchner
ef1fd341a8
EIGEN_STRONG_INLINE was NOT inlining in some critical needed areas (6.6X slowdown) when used with Tensorflow. Changing to EIGEN_ALWAYS_INLINE where appropiate.
2021-06-16 16:30:31 +00:00
Antonio Sanchez
9e94c59570
Add missing ppc pcmp_lt_or_nan<Packet8bf>
2021-06-15 13:42:17 -07:00
Rasmus Munk Larsen
fc87e2cbaa
Use bit_cast to create -0.0 for floating point types to avoid compiler optimization changing sign with --ffast-math enabled.
2021-06-11 02:35:53 +00:00
Antonio Sanchez
dba753a986
Add missing NEON ptranspose implementations.
...
Unified implementation using only `vzip`.
2021-05-25 18:25:35 +00:00
guoqiangqi
3d9051ea84
Changing the storage of the SSE complex packets to that of the wrapper. This should fix #2242 .
2021-05-10 23:53:16 +00:00
Christoph Hertzberg
722ca0b665
Revert addition of unused paddsub<Packet2cf>. This fixes #2242
2021-05-06 18:36:47 +02:00
Antonio Sanchez
1c013be2cc
Better CUDA complex division.
...
The original produced NaNs when dividing 0/b for subnormal b.
The `complex_divide_stable` was changed to use the more common
Smith's algorithm.
2021-04-29 17:39:58 +00:00
Antonio Sanchez
172db7bfc3
Add missing pcmp_lt_or_nan for NEON Packet4bf.
2021-04-27 14:12:11 -07:00
Jakub Lichman
d87648a6be
Tests added and AVX512 bug fixed for pcmp_lt_or_nan
2021-04-25 20:58:56 +00:00
Chip-Kerchner
06c2760bd1
Fix taking address of rvalue compiler issue with TensorFlow (plus other warnings).
2021-04-21 00:47:13 +00:00
Jakub Lichman
2b1dfd1ba0
HasExp added for AVX512 Packet8d
2021-04-20 19:07:58 +00:00
Antonio Sanchez
1d79c68ba0
Fix ldexp for AVX512 ( #2215 )
...
Wrong shuffle was used. Need to interleave low/high halves with a
`permute` instruction.
Fixes #2215 .
2021-04-20 16:25:22 +00:00
Christoph Hertzberg
9357feedc7
Avoid using uninitialized inputs and if available, use slightly more efficient movsd instruction for pset1<Packet2cf>.
2021-04-13 01:36:59 +02:00
Chip Kerchner
c24bee6120
Fix address of temporary object errors in clang11.
...
This fixes the problem with taking the address of temporary objects which clang11 treats as errors.
2021-04-02 16:27:08 +00:00
Antonio Sanchez
87729ea39f
Eliminate round_impl double-promotion warnings for c++03.
2021-03-25 16:52:19 +00:00
Chip Kerchner
d59ef212e1
Fixed performance issues for complex VSX and P10 MMA in gebp_kernel (level 3).
2021-03-25 11:08:19 +00:00
Christoph Hertzberg
69a4f70956
Revert "Uses _mm512_abs_pd for Packet8d pabs"
...
This reverts commit f019b97aca
2021-03-23 18:52:19 +00:00
David Tellenbach
4811e81966
Remove yet another comma at end of enum
2021-03-18 23:30:00 +01:00
Steve Bronder
f019b97aca
Uses _mm512_abs_pd for Packet8d pabs
2021-03-18 15:47:52 +00:00
Antonio Sanchez
8dfe1029a5
Augment NumTraits with min/max_exponent() again.
...
Replace usage of `std::numeric_limits<...>::min/max_exponent` in
codebase where possible. Also replaced some other `numeric_limits`
usages in affected tests with the `NumTraits` equivalent.
The previous MR !443 failed for c++03 due to lack of `constexpr`.
Because of this, we need to keep around the `std::numeric_limits`
version in enum expressions until the switch to c++11.
Fixes #2148
2021-03-16 20:12:46 -07:00
David Tellenbach
eb71e5db98
Fix another warning on missing commas
2021-03-17 03:07:04 +01:00
David Tellenbach
df4bc2731c
Revert "Augment NumTraits with min/max_exponent()."
...
This reverts commit 75ce9cd2a7 .
2021-03-17 03:06:08 +01:00
Antonio Sanchez
75ce9cd2a7
Augment NumTraits with min/max_exponent().
...
Replace usage of `std::numeric_limits<...>::min/max_exponent` in
codebase. Also replaced some other `numeric_limits` usages in
affected tests with the `NumTraits` equivalent.
Fixes #2148
2021-03-17 01:00:41 +00:00
David Tellenbach
9fb7062440
Silence warning on comma at end of enumerator list
2021-03-17 01:46:52 +01:00
Antonio Sanchez
f612df2736
Add fmod(half, half).
...
This is to support TensorFlow's `tf.math.floormod` for half.
2021-03-15 13:32:24 -07:00
Chip Kerchner
c9d4367fa4
Fix pround and add print
2021-03-15 19:07:43 +00:00
Antonio Sanchez
d24f9f9b55
Fix NVCC+ICC issues.
...
NVCC does not understand `__forceinline`, so we need to use `inline`
when compiling for GPU.
ICC specializes `std::complex` operators for `float` and `double`
by default, which cannot be used on device and conflict with Eigen's
workaround in CUDA/Complex.h. This can be prevented by defining
`_OVERRIDE_COMPLEX_SPECIALIZATION_` before including `<complex>`.
Added this define to the tests and to `Eigen/Core`, but this will
not work if the user includes `<complex>` before `<Eigen/Core>`.
ICC also seems to generate a duplicate `Map` symbol in
`PlainObjectBase`:
```
error: "Map" has already been declared in the current scope
static ConstMapType Map(const Scalar *data)
```
I tracked this down to `friend class Eigen::Map`. Putting the `friend`
statements at the bottom of the class seems to resolve this issue.
Fixes #2180
2021-03-15 18:42:04 +00:00
Antonio Sanchez
14487ed14e
Add increment/decrement operators to Eigen::half.
...
This is for consistency with bfloat16, and to support initialization
with `std::iota`.
2021-03-15 10:52:23 -07:00
Antonio Sanchez
853a5c4b84
Fix ambiguous call to CUDA __half constructor.
2021-03-08 21:06:28 -08:00
Antonio Sanchez
94327dbfba
Fix typo: DEVICE -> GPU
2021-03-08 11:21:00 -08:00
Antonio Sanchez
1296abdf82
Fix non-trivial Half constructor for CUDA.
...
Both CUDA and HIP require trivial default constructors for types used
in shared memory. Otherwise failing with
```
error: initialization is not supported for __shared__ variables.
```
2021-03-08 07:32:54 -08:00