11.1 The SSE/AVX Architectures 11.2 Streaming Data Types 11.3 Using cpuid to Differentiate Instruction Sets 11.4 Full-Segment Syntax and Segment Alignment 11.5 SSE, AVX, and AVX2 Memory Operand Alignment 11.6 SIMD Data Movement Instructions11.6.1 The (v)movd and (v)movq Instructions11.6.2 The (v)movaps, (v)movapd, and (v)movdqa Instructions11.6.3 The (v)movups, (v)movupd, and (v)movdqu Instructions11.6.4 Performance of Aligned and Unaligned Moves11.6.5 The (v)movlps and (v)movlpd Instructions11.6.6 The movhps and movhpd Instructions11.6.7 The vmovhps and vmovhpd Instructions11.6.8 The movlhps and vmovlhps Instructions11.6.9 The movhlps and vmovhlps Instructions11.6.10 The (v)movshdup and (v)movsldup Instructions11.6.11 The (v)movddup Instruction11.6.12 The (v)lddqu Instruction11.6.13 Performance Issues and the SIMD Move Instructions11.6.14 Some Final Comments on the SIMD Move Instructions 11.7 The Shuffle and Unpack Instructions11.7.1 The (v)pshufb Instructions11.7.2 The (v)pshufd Instructions11.7.3 The (v)pshuflw and (v)pshufhw Instructions11.7.4 The shufps and shufpd Instructions11.7.5 The vshufps and vshufpd Instructions11.7.6 The (v)unpcklps, (v)unpckhps, (v)unpcklpd, and (v)unpckhpd Instructions11.7.7 The Integer Unpack Instructions11.7.8 The (v)pextrb, (v)pextrw, (v)pextrd, and (v)pextrq Instructions11.7.9 The (v)pinsrb, (v)pinsrw, (v)pinsrd, and (v)pinsrq Instructions11.7.10 The (v)extractps and (v)insertps Instructions 11.8 SIMD Arithmetic and Logical Operations 11.9 The SIMD Logical (Bitwise) Instructions11.9.1 The (v)ptest Instructions11.9.2 The Byte Shift Instructions11.9.3 The Bit Shift Instructions 11.10 The SIMD Integer Arithmetic Instructions11.10.1 SIMD Integer Addition11.10.2 Horizontal Additions11.10.3 Double-Word–Sized Horizontal Additions11.10.4 SIMD Integer Subtraction11.10.5 SIMD Integer Multiplication11.10.6 SIMD Integer Averages11.10.7 SIMD Integer Minimum and Maximum11.10.8 SIMD Integer Absolute Value11.10.9 SIMD Integer Sign Adjustment Instructions11.10.10 SIMD Integer Comparison Instructions11.10.11 Integer Conversions 11.11 SIMD Floating-Point Arithmetic Operations 11.12 SIMD Floating-Point Comparison Instructions11.12.1 SSE and AVX Comparisons11.12.2 Unordered vs. Ordered Comparisons11.12.3 Signaling and Quiet Comparisons11.12.4 Instruction Synonyms11.12.5 AVX Extended Comparisons11.12.6 Using SIMD Comparison Instructions11.12.7 The (v)movmskps, (v)movmskpd Instructions 11.13 Floating-Point Conversion Instructions 11.14 Aligning SIMD Memory Accesses 11.15 Aligning Word, Dword, and Qword Object Addresses 11.16 Filling an XMM Register with Several Copies of the Same Value 11.17 Loading Some Common Constants Into XMM and YMM Registers 11.18 Setting, Clearing, Inverting, and Testing a Single Bit in an SSE Register 11.19 Processing Two Vectors by Using a Single Incremented Index 11.20 Aligning Two Addresses to a Boundary 11.21 Working with Blocks of Data Whose Length Is Not a Multiple of the SSE/AVX Register Size 11.22 Dynamically Testing for a CPU Feature 11.23 The MASM Include Directive 11.24 And a Whole Lot More 11.25 For More Information 11.26 Test Yourself