Intel 8087: Exploring the Hybrid CORDIC Algorithm

Ken Shirriff and others have reverse-engineered the tangent instruction (FPTAN) of the 1980 Intel 8087 floating-point coprocessor, uncovering a hybrid algorithm that combined 16 CORDIC iterations with a Padé approximant to deliver results in roughly 90 microseconds—about 144 times faster than software emulation on the 8086.

Microcode Extraction and the Three-Phase FPTAN Routine

To uncover how the 8087 handled trigonometric functions, Shirriff decapped an 8087 chip, photographed the silicon die under a microscope, and transcribed the 1,648 micro-instructions stored inside the central microcode ROM. The FPTAN routine occupies addresses #1039 through #1136 on the die. Depending on the input argument’s exponent, the code branches accordingly: arguments with exponents of minus-64 or less return the angle itself, while exponents between minus-17 and minus-64 skip CORDIC entirely to run a rational approximation. For inputs with exponents between minus-1 and minus-16, the chip executes a three-phase sequence.

According to Shirriff’s analysis, the operation begins with a CORDIC pseudo-division loop that subtracts precomputed special angles ($arctan(2^{-n})$) and records a 16-bit decision vector in a dedicated shift register. Next, a Padé approximant—specifically the rational polynomial $3x / (3 – x^2)$—computes the tangent of the remaining tiny residual angle of approximately $2^{-16}$. Finally, a CORDIC pseudo-multiplication loop applies those recorded rotations in reverse order to yield the final X and Y coordinates. Instead of performing a direct division, the instruction returns these as a numerator-denominator pair.

Hardware Architecture of the 8087 Datapath

The physical layout of the 8087 die mirrors its mathematical pipeline. The lower half of the silicon houses the datapath, containing specialized functional units that execute these floating-point calculations on 80-bit values. An exponent ROM stores fixed values required by the math routines, while a separate constant ROM holds the precomputed CORDIC constants. A large hardware shifter handles arbitrary 64-bit shifts left or right.

The heart of the arithmetic is a central adder which, alongside basic addition and subtraction, operates in a loop to handle multiplication, division, and square roots. Supporting this datapath are the B register, a sum register, eight stack registers, temporary registers, and a specialized shift register that retains the 16 status bits needed for CORDIC tracking.

Historical Origins of CORDIC and Intel Performance Metrics

The CORDIC (COordinate Rotation DIgital Computer) method dates back to 1956, when engineer Jack Volder developed it for the navigation computer of the Convair B-58 Hustler, the first bomber capable of flying at Mach 2. Because analog components provided limited accuracy and 1950s transistors were slow, Volder designed a digital algorithm that rotated vectors using only simple additions, subtractions, bit shifts, and table lookups, completely avoiding hardware multiplication or division.

For a typical input value like 0.95 radians, Shirriff’s cycle measurements show that the 8087 spends 33% of its time on CORDIC pseudo-division, 47% on CORDIC pseudo-multiplication, and just 15% on the rational approximation—dominated by a 64-bit squaring operation utilizing Booth’s radix-4 algorithm—with 5% overhead. While Intel’s official documentation cites a typical execution time of 450 clock cycles within a 30 to 540 cycle range, this hybrid design achieved a massive leap over the 13,000 microseconds required by software emulation on the accompanying 8086 processor.

Evolution Beyond the x87 Architecture

As semiconductor manufacturing advanced, Intel’s design philosophy shifted. By the Pentium series of CPUs, faster hardware multipliers made pure polynomial approximations practical, rendering CORDIC unnecessary since it struggled to scale to high bit counts without incurring major time penalties. The introduction of SIMD instructions and modern mathematical libraries like MKL and SVML eventually reduced the functional footprint of the x87 instruction set architecture and its legacy 80-bit temporaries.

Intel 8087: Exploring the Hybrid CORDIC Algorithm
Photo: righto.com

Hardware Context Note

Shirriff collaborated with members of the Opcode Collective, including Smartest Blob and Gloriouscow, on the microcode extraction. Annotated microcode listings and high-resolution die photographs are available via the official 8087 GitHub repository.

Purpose and mathematical algorithms of the Intel 8087

What was the primary purpose of the Intel 8087?

Introduced in 1980, the Intel 8087 was a floating-point coprocessor designed to accelerate mathematical calculations in the IBM PC and other host systems, providing massive speedups over software emulation on the 8086 microprocessor.

How the 8087 Outsmarted CORDIC for Tangents

Why did the 8087 use a hybrid algorithm for tangents?

The chip combined 16 CORDIC iterations with a Padé approximant to balance hardware simplicity with calculation speed and precision. CORDIC avoided complex multiplication hardware, while the polynomial stage compensated for CORDIC’s linear convergence near asymptotic boundaries like $pi/2$.

How did Ken Shirriff analyze the chip?

Shirriff decapped an 8087 coprocessor, photographed the silicon die under a microscope, and transcribed the 1,648 micro-instructions stored in the microcode ROM to map out the exact execution flow of instructions like FPTAN.

Leave a Comment