High Pass Image Filter
A multithreaded image-sharpening tool comparing C++ and x64 Assembly implementations of the same convolution filter.
Gallery
Overview
High Pass Image Filter is a solo university project I built to answer one question with real numbers: how much faster is hand-written x64 Assembly than C++ for the same image filter? A C# host application handles image loading, threading, benchmarking and saving, and calls into two interchangeable native DLLs that expose an identical C ABI, one written in C++ and one in MASM Assembly, so I could swap implementations at runtime and measure them head to head. The filter itself is a 3x3 high-pass convolution that sharpens edges and lifts local contrast.
Technical Highlights
- SIMD sum with
psadbw. The Assembly path packs the eight neighbouring pixels into an XMM register and usespsadbw(sum of absolute differences against zero) to add them in a single instruction instead of eight scalar operations, exploiting the kernel’s uniform-1coefficients. Code inAsm.asm. - Branchless clamping. The result is clamped to the valid 0-255 range with
cmovg/cmovlconditional moves rather than branches, keeping the per-pixel hot path free of misprediction stalls, also inAsm.asm. - Runtime-swappable native DLLs. Both implementations export the same C ABI (
extern "C" __declspec(dllexport)), so the host chooses either one at runtime through P/Invoke with no duplicated calling code. Seedllmain.cpp. - Row-range threading. The host splits the image into row ranges across a configurable thread count (1 to 64) with correct byte-offset math and remainder handling, preserving memory locality. See
ThreadsManager.cs. - Bulk bitmap I/O. Pixels are lifted into a flat byte array in one
Marshal.CopyviaLockBits, avoiding the per-pixelGetPixel/SetPixelpenalty. SeeCustomBitmap.cs. - Honest benchmarking. The harness trims the first 2.5% of samples to discard JIT warm-up, then reports average time with standard deviation across seven thread counts. See
TimeMesurement.cs.
The payoff: hand-written Assembly ran nearly twice as fast as the compiler’s best C++.
Learnings
This was my first serious Assembly work, and the biggest lesson was to measure rather than assume. Seeing exactly where SIMD and branchless code moved the numbers taught me more about the gap between source code and the CPU than any amount of reading.