JustinTime Compilation for FPGA Processor Cores Andrew Becker
Just-in-Time Compilation for FPGA Processor Cores Andrew Becker 1, Scott Sirowy 2, Frank Vahid Department of Computer Science and Engineering University of California, Riverside {abecker | ssirowy | vahid}@cs. ucr. edu 1. Now at EPFL 2. Now at ESRI This work was supported in part by the National Science Foundation (CNS 1016792) and by the Semiconductor Research Corporation (GRC 2143. 001)
Motivation System. C useful capture language Concurrency, structure, timing Simulation typical, but in-system I/O often useful Design/synthesis to FPGA may take hours/days and require advanced tools Switches/LEDs Simulation Cameras/displays In-system I/O
Background Want rapid design iteration with in-system I/O Compile design description; avoid design/synthesis Previously: Hybrid approach—System. C bytecode System. C Code class CLK_GEN : public sc_module { sc_in<bool> clock; … CLK_GEN(){ … Simulator (no in-system I/O) Design/synthesis (time-consuming) Compiler Portable System. C-on-a-chip – Sirowy [CODES+ISSS ’ 09] Bytecode process(clock) READ $1 data. Rdy BGT $1 $0 Start J Done Start: ADDI $2 $2 1 ADDI $3 $0 7 … …
Background Emulate bytecode in engine on FPGA class CLK_GEN : public sc_module { sc_in<bo ol> clock; … CLK_GEN( ){ … Fast compilation Bytecode also portable (FPGA-device independent) FPGA Bytecode Compiler process(clock) READ $1 data. Rdy BGT $1 $0 Start J Done Start: ADDI $2 $2 1 ADDI $3 $0 7 … Emulation Engine In-system I/O Portable System. C-on-a-chip – Sirowy [CODES+ISSS ’ 09]
Emulation Engine Discrete event simulator C code on a processor (Currently Microblaze soft-core; could be hard-core) Support-circuits for architectural features, peripheral I/O Peripheral Bus UART LEDs Buttons Frame Buffer Processor Core Event Kernel Instruction Mem. Read Signal Memory Write Signal Memory
Caveat Emptor Emulation is slow On soft-core, is even slower than PC simulation Won't meet many real-time constraints
This work – Speed up emulator First analyzed emulator performance
Low-Hanging Fruit 69% of time spent emulating bytecode Two strategies to reduce Reduce each instruction’s emulation time Reduce instruction memory latency
First Step Reduce instruction emulation time • Optimize event kernel? Peripheral Bus UART LEDs Buttons Frame Buffer Processor Core Event Kernel Instruction Mem. Read Signal Memory Write Signal Memory
First Step Reduce instruction emulation time • Optimize event kernel? • Just-in-time (JIT) compile bytecode to native processor code, done transparently by event kernel Peripheral Bus UART LEDs Buttons Frame Buffer Processor Core Event Kernel Instruction Mem. Read Signal Memory Write Signal Memory
Just-in-Time Compilation of Bytecode Implemented System. C-bytecode to Microblaze JIT compiler 3 x speedup; still portable Tunable delay/jitter Still want more speed Emulation Engine Machine Code Machine Bytecode process(clock) READ $1 data. Rdy BGT $1 $0 Start J Done Start: ADDI $2 $2 1 ADDI $3 $0 7 … JIT IMM 0 x. DEAD LWI $11 $0 0 x. BEEF BGTI $11 Start BRAI Done Start: … Event Kernel
Further Improvement Reduce instruction memory latency Add dedicated small, fast memory for JIT code on a fast, local bus Unique JIT possibility due to FPGA configurability
Architecture Changes Peripheral Bus UART Local Memory Bus Processor Core LEDs Buttons Frame Buffer Instr. Mem. JIT Mem. Read Signal Memory Write Signal Memory Emulation Engine
Even Further Improvement 23% of time spent maintaining signal queue What can be done? • Optimize signal queue maintenance code?
Common Denominator FPGA offers configurability Engine designer can make tradeoffs Trade hardware resources for speed FPGA Extra Resources Emulation Engine
Common Denominator FPGA offers configurability Engine designer can make tradeoffs Trade hardware resources for speed Add another soft-core? FPGA Extra Resources Emulation Engine
Even Further Improvement 23% of time spent maintaining signal queue What can be done? • Optimize signal queue maintenance code? Offload job to coprocessor • Again, unique JIT option due to FPGA configurability
Architecture Changes Peripheral Bus UART Local Memory Bus Processor Core LEDs Buttons Frame Buffer Instr. Mem. Signal Queue JIT Mem. Read Signal Memory Write Signal Memory Emulation Memory Controller Emulation Engine
Experimental Results 19
Conclusions • Approach rapid design iteration with in-system I/O • Uses • Education (typically loose timing constraints) • System prototypes that can tolerate real-time slowdown (e. g. , slow frame rate) • Portable and flexible • Engine design sets speed, not compiler or CAD flow • This work: 15 x speedup via normal JIT (3 x) + FPGA-specific JIT (5 x) • But, still orders of magnitude slower than design/synthesis • Future work: Bytecode accelerators, JIT synthesis
- Slides: 20