~ $ cd research && cat README.md
Research
In 2003 every new processor had run old programs faster for twenty-five years, and it seemed it would last for ever. My thesis began from the hunch that it would not: that the desktop would go multi-core, and that the problem would stop being the hardware and become the programmer, who now had to write many coordinated lists of instructions instead of one, and mostly did not know how.
So the question of the whole PhD is one of usability. Supercomputers had been parallel for decades, and most of their users were not computer scientists. What they used, OpenMP, lets you take a serial program and annotate it, a line at a time, until it runs in parallel. Could that way of working be carried to the machines everyone was about to own, which did not share memory, and whose cores were not all the same?
Algorithms Acceleration of Pattern-Matching in Multi-Core Architectures, Universitat Rovira i Virgili, defended in Tarragona on 8 July 2011, cum laude, directed by Francesc Serratosa. Eight years, two universities, and an unusual spread for one thesis: a runtime, a simulator, a compiler, and an algorithm.
What is in it
- Tools for parallel machines — UPC and the Barcelona Supercomputing Center, 2003–2008. OpenMP on a chip with 128 hardware threads, written with IBM Research; OpenMP on a cluster with no shared memory; a simulator of a heterogeneous processor; and a compiler that turns annotated serial C into streams, in a European project.
- Streams by annotation — the stream compiler by the examples of its own poster: a serial C program becomes a pipeline of tasks by writing in its margin what goes in and what comes out, with what the prototype measured.
- Graph matching on a desktop — Universitat Rovira i Virgili, 2009–2011. Two computer-vision algorithms rewritten for CUDA and OpenMP on an eighteen-watt desktop, with the measurements: up to forty times faster, without changing the result by a bit.
- Loops into zones — the method of the thesis, step by step: two loop transformations and two annotations that give a serial algorithm the shape of a graphics card, without changing what it computes.
- One lock at a time — in between the two, a year inside a database engine, making its core concurrent: each technique with the speed-up it bought, including the right change that made everything twice as slow.
Since
Two experiments I ran later, and wrote up on Medium:
- Coverage, generated — 2023. A program that knows nothing of what the code is for wrote the tests, and reached 85% code coverage: why coverage helps the programmer, and says nothing to a manager.
- The language of the question — 2025. The same questions to the same AI in five languages, and answers that did not agree.
What it left
The thesis ends on a sentence I still use: desktop computers are indeed desktop supercomputers, not only by their performance, but also by their complexity. Its tools were released under the GPL, and its last slide argued that research software should be published with its sources, the way a paper is published with its proofs.
And it left a habit. Everything I have built since for other engineers — a platform, a test harness, a course — starts from the question this started from: not what the machine can do, but what the person in front of it can be expected to get right. A small case of it: the recipe for concurrency I wrote a consensus algorithm by, so that students could.
The thesis lists sixteen publications. The record: the thesis, at Dialnet, and what DBLP indexes.
Use ls to see the parts, or cat parallel-tools to read one here.
~/research $ ls
parallel-tools/Tools for parallel machinesUPC and the Barcelona Supercomputing Center, 2003–2008. OpenMP on a chip with 128 threads and on a cluster with no shared memory, a simulator of a heterogeneous processor, and a compiler that turns serial C into streams.streams-by-annotation/Streams by annotationA serial C program becomes a pipeline of tasks by writing in its margin what goes in and what comes out. The model, by the examples of its own poster, and what the prototype measured.graph-matching/Graph matching on a desktopUniversitat Rovira i Virgili, 2009–2011. Two computer-vision algorithms rewritten for CUDA and OpenMP on an eighteen-watt desktop: an hour and a quarter became under two minutes, without changing the result by a bit.loops-into-zones/Loops into zonesThe method of the thesis, step by step. Two loop transformations and two annotations turn a serial algorithm into one shaped like a graphics card, without changing what it computes.one-lock-at-a-time/One lock at a timeMaking a database engine's core concurrent, one technique at a time, each with its number — including the right change that made everything twice as slow.coverage-generated/Coverage, generated2023. A program that knows nothing of what the code is for wrote the tests, and reached 85% code coverage, more than most companies ask for. Why coverage helps the programmer, and says nothing to a manager.language-of-the-question/The language of the question2025. The same questions to the same AI in Catalan, Spanish, English, French and Dutch, and answers that did not agree. On dubbing, each language took its own side; on tourism, the most concrete answers were not in English.