Rasmus Munk Larsen
75bcd155c4
Vectorize tan(x)
...
libeigen/eigen!2086
Co-authored-by: Rasmus Munk Larsen <rmlarsen@google.com >
2025-12-02 21:53:10 +00:00
Rasmus Munk Larsen
b6fcddccfc
Get rid of pblend packet op.
...
There was only a single code path left in TensorEvaluator using pblend. We can replace that with a call to the more general TernarySelectOp and get rid of pblend entirely from Core.
Closes #2998
See merge request libeigen/eigen!2056
Co-authored-by: Rasmus Munk Larsen <rmlarsen@google.com >
2025-11-03 23:27:50 +00:00
Rasmus Munk Larsen
ce70a507c0
Enable more generic packet ops for double.
...
See merge request libeigen/eigen!2050
2025-10-30 19:14:43 +00:00
Rasmus Munk Larsen
97c7cc6200
Explicitly use the packet trait HasPow to control whether Pow is vectorized.
2025-07-18 21:51:42 +00:00
Antonio Sánchez
26616fe5b8
Fix VSX packetmath psin and pcast tests.
2025-06-27 04:08:20 +00:00
Rasmus Munk Larsen
33f5f59614
Vectorize cbrt for float and double.
2025-04-17 23:31:20 +00:00
Antonio Sánchez
d935916ac6
Add numext::fma and missing pmadd implementations.
2025-03-23 01:05:53 +00:00
C. Antonio Sanchez
1d8b82b074
Fix power builds for no VSX and no POWER8.
2025-02-15 13:56:47 -08:00
Antonio Sanchez
74264c391a
Add missing return statements for ppc.
2025-02-05 08:12:27 -08:00
Antonio Sánchez
b1e74b1ccd
Fix all the doxygen warnings.
2025-02-01 00:00:31 +00:00
Rasmus Munk Larsen
5133c836c0
Vectorize erf(x) for double.
2024-11-16 19:05:16 +00:00
Rasmus Munk Larsen
0d366f6532
Vectorize erfc(x) for double and improve erfc(x) for float.
2024-11-08 17:21:11 +00:00
Rasmus Munk Larsen
58b252e5b3
Fix typo in PacketMath.h
2024-10-28 18:19:52 +00:00
Rasmus Munk Larsen
6c04d0cd68
Add missing exp2 definition for Altivec.
2024-10-28 18:12:36 +00:00
Rasmus Munk Larsen
bbdabebf44
Vectorize atanh<double>. Make atanh(x) standard compliant for |x| >= 1.
2024-08-30 17:27:55 +00:00
Rasmus Munk Larsen
32d95bb097
Add vectorized implementation of tanh<double>
2024-08-21 02:29:45 +00:00
Frédéric Chapoton
6331da95eb
fixing a lot of typos
2024-07-30 22:15:49 +00:00
Chip Kerchner
4d1d14e069
Change predux on PowerPC for Packet4i to NOT saturate the sum of the elements (like other architectures).
2024-05-08 22:39:27 +00:00
Charles Schlosser
fb95e90f7f
Add truncation op
2024-04-29 23:45:49 +00:00
Antonio Sanchez
1c8c734c8b
Fix sin/cos on PPC.
2024-04-24 15:58:03 -07:00
Antonio Sánchez
f0795d35e3
Fix new psincos for ppc and arm32.
2024-04-19 00:31:09 +00:00
Chip Kerchner
ad452e575d
Fix compilation problems with PacketI on PowerPC.
2024-04-18 14:55:15 +00:00
Chip Kerchner
be54cc8ded
Fix preverse for PowerPC.
2024-04-03 20:09:06 +00:00
Damiano Franzò
be06c9ad51
Implement float pexp_complex
2024-02-17 00:26:57 +00:00
Antonio Sánchez
3ebaab8a63
Fix PPC rand and other failures.
2024-02-05 20:07:15 +00:00
Damiano Franzò
7fd7a3f946
Implement plog_complex
2024-01-30 19:06:05 +00:00
Tobias Wood
f38e16c193
Apply clang-format
2023-11-29 11:12:48 +00:00
Chip Kerchner
4e598ad259
New panel modes for GEMM MMA (real & complex).
2023-09-06 20:03:45 +00:00
Antonio Sánchez
6e4d5d4832
Add IWYU private pragmas to internal headers.
2023-08-21 16:25:22 +00:00
Chip Kerchner
7769eb1b2e
Fix problems with recent changes and Tensorflow in Power
2023-07-26 16:24:58 +00:00
Marcus Comstedt
8f927fb52e
Altivec: fix compilation with C++20 and higher
2023-07-05 13:14:02 +00:00
Chip Kerchner
3791ac8a1a
Fix supportsMMA to obey EIGEN_ALTIVEC_MMA_DYNAMIC_DISPATCH compilation flag and compiler support.
2023-06-28 17:57:21 +00:00
Chip Kerchner
211c5dfc67
Add optional offset parameter to ploadu_partial and pstoreu_partial
2023-06-23 19:53:05 +00:00
Chip Kerchner
b8208b363c
Specialized loadColData correctly - fix previous BF16 GEMV MR
2023-05-04 16:38:17 +00:00
Chip Kerchner
fda1373a15
Fix ColMajor BF16 GEMV for when vector is RowMajor
2023-05-03 20:12:50 +00:00
Chip Kerchner
6418ac0285
Unroll F32 to BF16 loop - 1.8X faster conversions for LLVM. Use vector pairs for GCC.
2023-05-01 16:54:16 +00:00
Chip Kerchner
03f646b7e3
New VSX version of BF16 GEMV (Power) - up to 6.7X faster
2023-04-21 17:06:59 +00:00
Chip Kerchner
3f3ce214e6
New BF16 pcast functions and move type casting to TypeCasting.h
2023-04-18 02:38:38 +00:00
Chip Kerchner
1148f0a9ec
Add dynamic dispatch to BF16 GEMM (Power) and new VSX version
2023-04-14 22:20:42 +00:00
Rasmus Munk Larsen
df1049ddf4
Small packet math cleanup.
2023-04-04 16:14:32 +00:00
Chip Kerchner
d71ac6a755
Fix recent PowerPC warnings and clang warning
2023-03-15 16:50:46 +00:00
Chip Kerchner
23e1541863
Put deadcode checks back in from previous change.
2023-03-14 00:57:16 +00:00
Chip Kerchner
6c58f0fe1f
Revert changes that made BF16 GEMM to cause bad register spillage for LLVM (Power)
2023-03-13 23:36:06 +00:00
Chip Kerchner
9d72412385
Add MMA to BF16 GEMV - 5.0-6.3X faster (for Power)
2023-03-13 19:37:13 +00:00
Rasmus Munk Larsen
ee0ff0ab3a
Fix typo in MathFunctions.h
2023-03-13 15:50:40 +00:00
Rasmus Munk Larsen
d6235d76db
Clean up generic packetmath specializations for various backends with the help of a macro.
2023-03-10 22:02:23 +00:00
Chip Kerchner
2b513ca2a0
Added partial linear access for LHS & Output - 30% faster for bfloat16 GEMM MMA (Power)
2023-03-02 19:22:43 +00:00
Chip Kerchner
e4598fedbe
Fix compiler versions for certain instructions on Power.
2023-02-23 23:24:41 +00:00
Rasmus Munk Larsen
ce62177b5b
Vectorize atanh & add a missing definition and unit test for atan.
2023-02-21 03:14:05 +00:00
Chip Kerchner
e797974689
Add and enable Packet int divide for Power10.
2023-02-17 19:04:18 +00:00