derek/gem5 - gem5 - Gitea: Git with a cup of tea

derek/gem5

Author	SHA1	Message	Date
Harshil Patel	47c4dad869	arch-riscv: Remove unnecessary assert (#866 ) `assert(interruptID >=0)` is always true as `interruptID` is an unsigned int. This was causing compilation tests failures in GCC-8 with the following error: ```sh src/arch/riscv/interrupts.cc:47:32: error: comparison is always true due to limited range of data type [-Werror=type-limits] assert(interruptID >= 0); ``` Change-Id: I356be78d7f75ea5d20d34768fb8ece0f746be2fc	2024-02-13 08:30:18 -08:00
Vishnu Ramadas	8054459df6	arch-vega: Add support for S_ICACHE_INV instruction Previously, the S_ICACHE_INV instruction was unimplemented and simulation panicked if it was encountered. This commit adds support for executing the instruction by injecting a memory barrier in the scalar pipeline and invalidating the ICACHE (or SQC) Change-Id: I0fbd4e53f630a267971a23cea6f17d4fef403d15	2024-02-09 12:19:08 -06:00
Saúl	7d80658a39	arch-riscv: fix vl in mask load/store (i.e vlm.v/vsm.v) (#830 ) The vlm.v and vsm.v unit-stride mask load/store instructions are constructed with an incorrect VL when the current one is larger than than VLEN/EEW (i.e. when LMUL > 1). This commit fixes the issue for both instructions.	2024-02-08 14:06:49 -08:00
Bobby R. Bruce	7fe1588546	arch-riscv: Fix load and store to use EEW instead of SEW (#859 ) Vector unit-stride instructions have an EEW encoded directly in the instruction, We should use that instead of SEW in vtype. Ref: https://github.com/riscv/riscv-v-spec/blob/master/v-spec.adoc#73-vector-loadstore-width-encoding	2024-02-08 12:14:11 -08:00
Saúl	804f137325	arch-riscv: add unit-stride fault-only-first loads (i.e. vleff) (#794 ) This patch provides unit-stride fault-only-first loads (i.e. vleff) for the RISC-V architecture. They are implemented within the regular unit-stride load (i.e. vle). A snippet named `fault_code` is inserted with templating to change their behaviour to fault-only-first. A part from this, a new micro based on the vset\vl\* instructions (VlFFTrimVlMicroOp) is inserted as the last micro in the macro constructor to trim the VL to it's corresponding length based on the faulting index. This trimming micro waits for the load micros to finish (via data dependency) and has a reference to the other micros to check whether they faulted or not. The new VL is calculated with the VL of each micro, stopping on the first faulting one (if there's such a fault). I've tested this with VLEN=128,256,...,16384 and all the corresponding SEW+LMUL configurations. Change-Id: I7b937f6bcb396725461bba4912d2667f3b22f955	2024-02-08 09:15:58 -08:00
QQeg	e685c072d1	arch-riscv: Remove micro_elems in VleMicro template Change-Id: I91267de8b1142075aa2873bfcedfd8b15c6863d4	2024-02-08 07:24:55 +00:00
QQeg	7eeac98b8d	arch-riscv: Fix load and store to use EEW instead of SEW Vector unit-stride instructions have an EEW encoded directly in the instruction, We should use that instead of SEW in vtype. Change-Id: I282041ce8ed57fbcca899f7497ef6c6fb2dfcf85	2024-02-07 21:11:28 +00:00
Robert Hauser	f289f9e8b5	arch-riscv: adding support for local interrupts (#813 ) Besides the standard RISC-V interrupts software, timer, and external interrupt, the RISC-V specification also offers the possibility to implement local interrupts. With this patch, we contribute an extension of RiscvInterrupts that enables connecting interrupt sources to the local interrupt controller. We assigned the local interrupts to machine-level and gave them the highest priority. If two local interrupts are pending, there exception code will be the tie-breaker (higher ID > lower ID). 32 Bit systems only recognize the local interrupts 16 to 31, 64 Bit systems 16 to 63. Change-Id: Iff8d34e740b925dce351c0c6f54f4bd37a647e0c --------- Co-authored-by: Robert Hauser <robert.hauser@uni-rostock.de>	2024-02-06 09:38:50 -08:00
Yu-Cheng Chang	ba6c569b8d	arch-riscv: Add BasePMAChecker to support customized PMA (#846 ) The RISC-V privilege spec don't specify the implementation of PMA(physical memory attribute), which is addressed in the previous CL[1]. This CL creates the BasePMAChecker to support customized PMA so that we can only focus on the features wanted in the study. The CL also leaves the common methods `check` and `takeOverFrom` to make MMU easy to interact with PMA. [1] https://gem5-review.googlesource.com/c/public/gem5/+/40596 Change-Id: I9725e3a8f7f9276e41f0d06988259456149d2a77	2024-02-06 05:38:34 -08:00
Giacomo Travaglini	a60d6960c7	arch-arm: Remove unused/unimplemented TLB methods (#849 ) Change-Id: I3a76a914df1ba65ec5200f11111cf20f3e1eb924 Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-02-06 09:18:06 +00:00
Giacomo Travaglini	05f93175a7	arch-arm: Crypto instruction execution requires SIMD to be enabled (#848 ) Crypto instructions will cause an undefined instruction when executed with SIMD disabled. The PR is also refactoring their implementation by checking the release object instead of the ID register field. This is improving readability	2024-02-05 19:22:04 +00:00
Chong-Teng Wang	40ecdf5fb4	arch-riscv: Fix RVV instructions vmv.s.x/vfmv.s.f (#843 ) This commit fixes the implementation of vmv.s.x and vfmv.s.f. When vl = 0, no elements are updated in the destination vector register group, regardless of vstart. Change-Id: Ib21b3125da8009325743ec70ca0874704328356c Reference: [Integer Scalar Move Instructions](https://github.com/riscv/riscv-v-spec/blob/master/v-spec.adoc#161-integer-scalar-move-instructions) [Floating-Point Scalar Move Instructions](https://github.com/riscv/riscv-v-spec/blob/master/v-spec.adoc#162-floating-point-scalar-move-instructions)	2024-02-05 08:51:42 -08:00
Chong-Teng Wang	85059a369e	arch-riscv: Fix control flow in VectorFloatMaskMacroConstructor (#844 ) This commit adjusts the logic in VectorFloatMaskMacroConstructor to ensure the %(copy_old_vd)s section is not skipped when vl = 0, ensuring correct values in destination vector register. Change-Id: I2478722d6f003a0f2e4b3cd0ba3e845bed938ee6 This is the same problem as #715 .	2024-02-05 06:29:05 -08:00
Giacomo Travaglini	16e06bad0c	arch-arm: Exec Crypto instructions only if SIMD&FP enabled We not only check for the presence of the relative FEAT_*, we also check if AdvSIMD is enabled; we throw an undefined instruction otherwise. Change-Id: I1fd0cdc8057c5a7901774802dc076817f06c8e66 Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com> Reviewed-by: Andreas Sandberg <andreas.sandberg@arm.com>	2024-02-05 12:56:48 +00:00
Giacomo Travaglini	ebef2fc4b1	arch-arm: Crypto instructions checking release object Check directly if extension is enabled instead of looking for ID register field value. This makes the code more readable Change-Id: If0b882ac3464c3587731b72a7edb3b8b65ea86c7 Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com> Reviewed-by: Andreas Sandberg <andreas.sandberg@arm.com>	2024-02-05 12:56:48 +00:00
Giacomo Travaglini	3a2c8feca8	arch-arm: MMU aarch64EL is not used in AArch64 only anymore We therefore rename it to exceptionLevel Change-Id: I2a3aabaefa315d95bd034b13d95d5a5b0b8e9319 Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-02-01 13:45:06 +00:00
Giacomo Travaglini	3737e8b6df	arch-arm: Use MAIR_EL2 mem attribute register when in EL0 host With the old code, the MAIR_EL1 register was checked when inserting an EL2&0 TLB entry Change-Id: I064032fb2946777c2f4c50c06a124f828245e18a Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-02-01 13:44:16 +00:00
Giacomo Travaglini	d42ef792bf	arch-arm: Check ELIs64 for EL2 when in EL2&0 regime The problem with: ELIs64(tc, aarch64EL == EL0 ? EL1 : aarch64EL); Is that when we are executing at EL0 in host (EL2&0 translation regime), the execution mode (AArch32 vs AArch64) is dictated by EL2 and not by EL1 (which is the guest) Change-Id: I463a2a9461c94d0886990ae3d0a6e22aeb4b9ea3 Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-02-01 13:43:59 +00:00
Giacomo Travaglini	458c98082c	arch-arm: Replace EL based translation with regimes This is the final step in the transformation process. We limit the use of the "managing Exception Level" for a translation in favour of the more standard "Translation Regime" This greatly simplifies our code, especially with VHE where the managing el (EL2) could handle to different translation regimes (EL and EL2&0). We can therefore remove the isHost flag wherever it got used. That case is automatically handled by the proper regime value (EL2&0) Change-Id: Iafd1d2ce4757cfa6598656759694e5e7b05267ad Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-02-01 13:43:47 +00:00
Giacomo Travaglini	e333a77c12	arch-arm: Remove _Xt postfix from TLBI instructions The Xt is not part of the architectural name of the register and it was likely added with the introduction of extended register (Xt) TLBIs in Armv8 to differentiate them with the old Armv7 ones. The use of _Xt was not consistent anyway: newer TLBIs were already omitting it. Change-Id: Ic805340ffa7b5770e3b75a71bfb76e055e651f8b Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-02-01 13:43:26 +00:00
Giacomo Travaglini	594428f010	arch-arm: Remove redundant isHyp as a TLB entry field We should stop using isHyp.. An hypervisor entry is flagged already by the EL of the entry (el == EL2) Change-Id: I20c3d06fa2b04e0b938a380ca917d0b596eddcf2 Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-02-01 13:43:00 +00:00
Giacomo Travaglini	a6ca81906a	arch-arm: Simplify setting of isHyp for mem translations The isHyp descriptor is an old artifact of armv7 and it flags a PL2 (AArch32) or EL2 & EL2&0 (AArch64) translations. It is commonly set according to the EL/mode [1] but it may differ from the execution state in case of explicit translation requests (via the AT instruction as an example [2]). There is really no need to complicate the masking of isHyp. We should just make use of the tranType method (in charge of setting aarch64EL) to properly set aarch64EL, and make isHyp coincide with the case of aarch64EL == EL2. This is a step towards the removal of the isHyp flag. More specifically the patch does the following: * HypMode translation type moved in the EL2 case The translation is used by ATS1HR/ATS1HW: Performs stage 1 address translation as defined for PL2 and the Non-secure state * S1S2NsTran translation type moved in the EL1 case The translation is used by ATS12NSOPR/ATS12NSOPW: Performs stage 1 and 2 address translations as defined for PL1 and the Non-secure state * S1CTran translation type can be at either EL1 or EL3 The translation is used by ATS1CPR/ATS1CPW Performs stage 1 address translation as defined for PL1 and the current Security state [1]: https://github.com/gem5/gem5/blob/stable/src/arch/arm/mmu.cc#L1281 [2]: https://github.com/gem5/gem5/blob/stable/src/arch/arm/mmu.cc#L1282 Change-Id: Ie653170f6053c5d8141a2de9f50febf5bf53ab9c Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-02-01 13:42:40 +00:00
Jason Lowe-Power	b3870ee7b0	arch-riscv: Fix fence.i instruction in O3 CPU (#816 ) arch-riscv: Fix fence.i instruction in O3 CPU	2024-01-30 15:39:32 -08:00
Roger Chang	d94ef08a36	arch-riscv: Fix fence.i instruction in O3 CPU We should clean the instruction buffer after the fence.i is execute to avoid execute old instruction for self-modifying code Change-Id: Iece0ee0d10631fcd9bd17ee67cf0c92f72acdbd8	2024-01-29 11:43:27 +08:00
QQeg	08ed87bc9d	arch-riscv: Add template Vector1Vs1VdMaskDeclare This commit adds a new template, Vector1Vs1VdMaskDeclare, to replace the use of Vector1Vs1RdMaskDeclare in Vector1Vs1VdMaskFormat. The change addresses the issue with the number of indices in srcRegIdxArr. Only two indices are available in Vector1Vs1RdMaskDeclare, but instructions that use Vector1Vs1VdMaskFormat, like 'vmsbf', require three indices (for vs1, vs2(old_vd), and vm) to function correctly. Change-Id: I0c966e11289ce07efcc3b0cc56948311289530ad	2024-01-28 09:38:11 +00:00
QQeg	31ffc11c57	arch-riscv: Fix segmentation fault in vmsbf/vmsof/vmsif This commit simplifies the conditional logic in vmsbf/vmsof/vmsif by removing an unnecessary variable and condition. The updated logic checks 'this->vm' or the result of 'elem_mask(v0, i)' directly, which prevents a segmentation fault regardless of whether 'vm' is set or not. Change-Id: I799fa7b684ff98959a64f9694ef9c854f3a1f43a	2024-01-28 09:38:11 +00:00
Giacomo Travaglini	ce32d7c523	arch-arm: Replace CRYPTO extension with canonical names (#810 ) These are: FEAT_AES, FEAT_PMULL, FEAT_SHA256, FEAT_SHA1, FEAT_CRC32 With this patch we are also enabling them by default by adding them to the Armv8 release object. Some of them are mandatory anyway since Armv8.1 Change-Id: I221ae8646d91151fdfaf97a4815168a4fe3d8c5a Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-01-26 19:39:35 +00:00
Ivana Mitrovic	24e0d71034	arch-gcn3: Remove gcn3 (#781 ) Related to issue #703 , this PR removes GCN3 related files and updates source code, documentation, and tests to switch over to Vega is that was not done already. Highlights are: - Remove all src/arch/amdgpu/gcn3 files and update Kconfigs. - Remove references to GCN3 and replace with Vega where applicable. - Update the build targets in the gcn-gpu Docker. This will need to be rebuilt but not urgently. - Remove the GCN3 tag in testlib. Most tests seem to be using Vega already, so that commit is small.	2024-01-25 10:14:46 -08:00
QQeg	7a96709b11	arch-riscv: Fix vsadd_vi and vsaddu_vi to match v-spec (#805 ) This commit fixes the implementation of two instructions, vsadd_vi and vsaddu_vi, in the OPIVI category to match the RISC-V vector specification. According to [riscv-v-spec](https://github.com/riscv/riscv-v-spec/blob/master/v-spec.adoc#101-vector-arithmetic-instruction-encoding), the immediate field of these two instructions should be sign extended. > For integer operations, the scalar can be a 5-bit immediate, imm[4:0], encoded in the rs1 field. The value is sign-extended to SEW bits, unless otherwise specified. There is an example in both [vsadd](https://github.com/QQeg/rvv_intrinsic_testcases/tree/master/vsadd_vi) and [vsaddu](https://github.com/QQeg/rvv_intrinsic_testcases/tree/master/vsaddu_vi). Change-Id: Ib877627ba01c0868b2103d41613651df488fca13	2024-01-24 17:21:26 -08:00
Yu-Cheng Chang	6dd936e5b5	arch-riscv: Simply implementation of vector multiply and divide instructions (#793 ) Align the implementation of scalar multiply and divide instructions Change-Id: I53297d4c841c41593baaae0ea140bfbbd874a1d9	2024-01-24 13:20:15 -08:00
Matthew Poremba	44c78d843c	arch-vega: Implement memory aperture operands (#803 ) Vega (gfx900) introduced new memory aperture registers to get the base address and limit for LDS and private (scratch) memory. These have not commonly been used by the compiler until ROCm 6. Now that the compiler is generating reads from these special registers, implement the support for them. Tested with LULESH which is using the SHARED_BASE register (LDS) with ROCm 6.0. This assembly seems to replace S_GETREG_B32 emitted by the ROCm 5 compiler. Change-Id: Id2bd26ce8ef687c84a647fa2ac2da54d657913e5	2024-01-24 11:19:43 -08:00
Matthew Poremba	dfafc5792a	arch-vega: Remove deleted instruction.cc from build (#801 ) Change-Id: I03073d35a0d36788dfe8309e6ed466d0a496e31e	2024-01-23 18:47:01 -08:00
Matthew Poremba	a5757e7e01	arch-vega: Rename mismatched source/header files The files registers.cc, isa.cc, and decoder.cc do not match the header name. This is a minor cleanup to make development more straightforward. Change-Id: Ibab18dfe315b0ce84359939b490f8227ea43cac0	2024-01-19 13:32:24 -06:00
Matthew Poremba	cd91c6321f	arch-vega: Reorganize instructions to multiple files The Vega instructions.cc file is 47k lines long which results in both large compilation times whenever it is modified and long style check times. This makes iterating over more complex instruction implementations very time consuming. This commit moves the instruction definitions to multiple files based on the instruction encoding (SOP2, VOP2, FLAT, DS, etc.). The resulting files are much smaller (max is 8k lines) and compilation and style check times are much more reasonable. Other than moving code around, there are no functional changes in this commit. Change-Id: Id4ac8e98ef11a58de5fd328f8a0cd7ce60a11819	2024-01-19 13:32:24 -06:00
Jason Lowe-Power	a555449c12	arch-arm: Fix compile error in kvm (#784 ) The addition of std::optional in #732 caused a compile error. This change fixes the error by checking to see if the value is present and panicing otherwise. Change-Id: I46c3fb76eb0e14ba7bede7c336293fbe9add8c84 Signed-off-by: Jason Lowe-Power <jason@lowepower.com>	2024-01-19 07:59:59 -08:00
Yu-Cheng Chang	f56459470a	arch-riscv: Refactor the RISC-V multiplication utility (#780 ) 1. Add the new double width for int64_t and uint64_t 2. Use the wider type to get the upper result of multiplication Change-Id: Id6cfa6f274c65592b2b3e2b70c00f82954b41f1a	2024-01-18 12:40:11 -08:00
QQeg	511729ab76	arch-riscv: Fix issue when vl=0 in VectorIntMaskMacroConstructor (#715 ) I’ve been working on a fix for the issue #759 where ‘vd’ incorrectly stores all zeros when ‘vl’ is set to 0 in VectorIntMaskMacroConstructor. My solution seems to work, but it behaves differently from other macros when ‘vl’ = 0. Instead of pushing a ‘nop’ to ‘microops’, it pushes a micro operation that remains ineffective due to ‘vl’ being 0.	2024-01-17 08:45:08 -08:00
Matthew Poremba	57fb083f43	arch-gcn3: Remove all GCN3 files Change-Id: Ib7d9e8676a31e51a330e68d81099580e2509a90a	2024-01-17 10:44:44 -06:00
Matthew Poremba	70376d43a3	arch-vega: Fix upsize cast error in newer compilers (#774 ) Newer compilers error on -Warray-length in the recent MI200 patches due to casting from a 32-bit data type to a 64-bit type. Change it to cast the 32-bit integer first then 64-bit integer latter to remove the warning. Rerun of validation tests on the three instructions passed. Change-Id: I0309e5f7b5b8cc8ce1651660ddddb120fa6e7666	2024-01-16 09:41:23 -08:00
Matthew Poremba	6a9e80c54c	gpu-compute: Support for MI200 GPU model (#733 )	2024-01-15 08:18:34 -08:00
Hoa Nguyen	85eb99388a	arch-riscv: Remove the check of bit 63 of the physical address (#756 ) Currently, the TLB enforces that the bit 63 of a physical address to be zero. This check stems from the riscv-tests that checks for the bit 63 of a physical address [1]. This is due to the fact that the ISA implicitly says that the physical address must be zero-extended on the most significant bits that are not translated [2]. More details on this issue is here [3]. The check for bit 63 of a physical address in the TLB is rather too specific, and I believe the check of invalid physical address is alread implemented in PMA. Thus, this change proposes to remove this check from RISC-V TLB. [1] `bd0a19c136/isa/rv64mi/access.S (L18)` [2] https://groups.google.com/a/groups.riscv.org/g/isa-dev/c/8kO7X0y4ubo [3] https://github.com/gem5/gem5/issues/238 Change-Id: I247e4d4c75c1ef49a16882c431095f6e83f30383 Signed-off-by: Hoa Nguyen <hn@hnpl.org>	2024-01-12 15:17:49 -08:00
Yu-Cheng Chang	2f24ee570e	arch-riscv: Move PMAChecker and PMP to RiscvISA namespace (#691 ) The PMAChecker and PMP are only used in the RisvISA and it should be in the RiscvISA to simply the implementation Change-Id: I4968e2de4c028cb2dceed977f2173fc8b1efd175	2024-01-10 16:58:13 -08:00
Yu-Cheng Chang	74dd0bb9bb	fastmodel: Fix the Fastmodel RemoteGDB initial (#735 ) Change-Id: Iec9ef145ccac353b8a41f501dd76bf53288dd478	2024-01-10 16:55:54 -08:00
Giacomo Travaglini	5e2e748f3a	arch-arm: Handle invalid case for encodeAArch64SysReg (#732 ) This patch is amending encodeAArch64SysReg so that it covers the case where there are no arch numbers available for the misc index passed as an argument. This could happen if the register ID is a gem5 pseudo register which is not associated with any architected op1/op2/crn/crm tuple. Rather than panicking we return a nullopt. Change-Id: I7ab70467105ef93c0c78ac4e999c7dc8e5e09925 Signed-off-by: Giacomo Travaglini <giacomo.travaglini@arm.com>	2024-01-04 10:04:40 +00:00
Matthew Poremba	31e63b01ad	arch-vega: Add vop3p DOT instructions Implemented according to the ISA spec. Validated with silion. In particular the sign extend is important for the signed variants and the unsigned variants seem to overflow lanes (hence why there is no mask() in the unsigned varints. FP16 -> FP32 continues using ARM's fplib. Tested vs. an MI210. Clamp has not been verified. Change-Id: Ifc09aecbc1ef2c92a5524a43ca529983018a6d59	2024-01-03 15:41:06 -06:00
Matthew Poremba	420cda1bef	arch-vega: Implement FP32 packed math Starting with MI200, packed math can operate on double dword inputs. In this case, 64-bits of inputs (two VGPRs per lane) contain two FP32 values. Add instructions to perform add, multiply, and FMA on packed FP32 types. Change-Id: Ib838bff91a10e02e013cc7c33ec3d91ff08647b0	2024-01-03 15:41:06 -06:00
Matthew Poremba	7b0c47d52f	arch-vega: Implement all global atomics up to gfx90a This change adds all of the missing flat/global atomics up to including the new atomics in gfx90a (MI200). Adds all decodings and instruction implementations with the exception of __half2 which does not have a corresponding data type in gem5. This refactors the execute() and completeAcc() methods by creating helper functions similar to what initiateAcc() uses. This reduces redundant code for global atomic instruction implementations. Validated all except PK_ADD_F16, ADD_F32, and ADD_F64 which will be done shortly. Verified the source/dest register sizes in the header are correct and the template parameters for the new execute()/completeAcc() methods are correct. Change-Id: I4b3351229af401a1a4cbfb97166801aac67b74e4	2024-01-03 15:41:06 -06:00
Matthew Poremba	472c697d88	arch-vega: Implement v_mfma_i32_16x16x16i8 Tested using AMD labs notes examples located on github: https://github.com/amd/amd-lab-notes/blob/release/matrix-cores/ src/mfma_i32_16x16x16i8.cpp Change-Id: Ib0e50162288528012b6d3395e1f629ebf12e8e54	2024-01-03 15:41:06 -06:00
Matthew Poremba	7e1b27969f	arch-vega: Improve FLAT disassembly Use the opSelectorToRegSym which will print the full range of VGPRs (e.g., will now print v[2:3] instead of v2 when the source / dest is 64-bits). Fixes atomic disassembly prints. Now shows "glc" if GLC bit is enabled. Fixes some VGPR fields being printed as an SGPR in places where the 9-bit register index bit is implied (e.g., VDST). This makes it easier to use a GPUExec trace to match with LLVM disassembly when debugging. Change-Id: Ia163774850f0054243907aca8fc8d0361e37fdd5	2024-01-03 10:40:34 -06:00
Matthew Poremba	bc69ab0a1f	arch-vega: Add VOP3P encodings and packed 16b insts This adds the VOP3P and VOP3P_MAI encodings from the MI200 spec. These instructions are used for packed math and miSIMD instructions. The first 19 VOP3P opcodes are implemented and validated against hardware. This includes all instructions which operate on one dword containing two packed 16-bit values of fp16, int16_t, or uint16_t. Implement one MFMA instruction for now which was also validated against hardware.	2024-01-03 10:40:34 -06:00

1 2 3 4 5 ...

5824 Commits