garganzol
4 days ago
100x speed improvement of math operations by 8087 is not an overestimation. The difference for apps relying on math was crazy back then. I experienced this first-hand on my 80286 machine, where it was 3-second vs 300-second calculation results.
One neat feature of 8087 instruction set is that it can be interspersed with x86 instructions in the code stream, giving you a simultaneous access to two processor chips working in parallel. This combo forms a real asymmetrical multi-processor system with certain opportunities for hardware-assisted code parallelization. If a thoughtful instruction scheduling is used, floating operations executed by 8087 work in parallel with the usual integer x86 code.
bell-cot
3 days ago
> 100x speed improvement ...
Vs. Ken's Blog says "up to 100 times", and Wikipedia gives a lower estimate.
Theory: Your 100X experience compared Intel's "exact emulation" code (noted in the article) with native x87. That emulation would have to cover the myriad x87 oddities and corner cases which Ken describes. Vs. Ken's & Wikipedia's are comparing x87 to various "good enough" 8088 floating point libraries - so naturally much faster than Intel's exact code.
(And yes, speed might have been a low priority for the team writing Intel's emulator.)
mgaunard
4 days ago
Any modern processor has different execution ports specialized in different things and replicated a different number of times, and all of them can execute instructions in parallel.
It schedules to these transparently for you, that's known as superscalar execution. To maximize occupation, out-of-order execution and simultaneous multithreading are used.
Sharlin
4 days ago
Sure, but this was five (or six?) generations before actual superscalar x86 processors.
vlovich123
4 days ago
Three - after 8086 you had 80286, 80386, 80486 and then Pentium (superscalar).
dpq
4 days ago
80186 is often forgotten to have existed because it didn't see much success in the market / because IBM skipped it and went with 80286 for the AT, but it did exist.
vlovich123
4 days ago
The 80186 didn’t really introduce anything new architecturally. It’s basically an 8086 with a few more chips bundled on-die. 80286 however introduced protected mode, expanded the address size to 24 bit, hardware enforce memory protection, multitasking, etc.
hyperman1
3 days ago
80186 Was not IBM PC compatible. It had e.g. a PIC, DMA and timer built-in, and these were incompatible with the chips in a 8086 based IBM PC.
There were a few new instructions, too, mostly closing holes. You could left or right shift with a constant, while the 8086 had only 1 or the CX register. I think mul also gained a constant.
The fact that Intel released a CPU that could not be put in a PC probably indicates how low they estimated the survivability of the PC.
icedchai
3 days ago
There were "PC compatibles" that used it, like the Tandy 2000. You are correct in that they may not have been entirely compatible, but they did exist.
hyperman1
2 days ago
I had a tandy of that era, with i think 64K ram. The incompatibilities were big enough that almost nothing could run.
pkaye
4 days ago
I believe 80186 was for embedded systems.
librasteve
3 days ago
nope
Sharlin
4 days ago
Four because you have to count 8086 itself (486 was one, not zero, geverations before the Pentium and so on).
fulafel
4 days ago
Transparent scheduling of superscalar execution was a later advance in microprocessors, termed out of order execution. Apart from micros both did come out around the same time in the mid-1960s.
In the x86 microarchitectures superscalar came in Pentium and OoO got introduced in Pentium Pro.
(Superscalar is just having >1 pipelines, which at its introduction meant needing to manually schedule your code very carefully to take advantage of it absent the OoO execution. For example the frequently posted-about Doom optimizations and talk of the u and v pipes are about this. The scheduling didn't happen transparently in early superscalars, at best the cpu automatically stalled, and some archs (eg MIPS, i860, TI C3x) even visibly punted hardware detection of pipeline hazards and required the code to just not go there, see load delay slots and branch delay slots. )
mgaunard
2 days ago
It's already transparent even without out-of-order execution, and pipelining is an entirely different thing than superscalar execution.
fulafel
2 days ago
I'd argue most people would consider it scheduling only if there's some attempt to arrange order of execution to improve throughput. Idling the other execution units when they could be executing instructions is not scheduling, at least not good scheduling.
In estabilished computer architecture terminology, scheduling is decidedly a OoO execution term. (At least if you're talking about processors. There's also static scheduling, which means the compiler does it and the hardware doesn't.)
mgaunard
6 hours ago
Then most people would be wrong.
A VLIW ISA is a classical example of a superscalar processor without out-of-order execution.
rwmj
3 days ago
I remember getting a 80387 (coprocessor for the 80386) and POVRAY renders going from running for days to still many minutes but you could sit and watch it.