CENTRALE LYON - Post-doctoral researcher (12 months) Mathematics of energy-efficient AI: diffusion models on emerging hardware
ecolecentraledelyon · Ecully, Auvergne-Rhône-Alpes, France
About The Role
The question
Generating one image with a diffusion model means evaluating a large neural network tens to hundreds of times. That repetition is where the energy goes, and it is part of why data-centre demand is now visible at grid scale: the Lawrence Berkeley National Laboratory puts US data-centre electricity use at 176 TWh in 2023, 4.4 % of national consumption, a figure that more than doubled since 2017 largely because of AI servers, and projects 325–580 TWh by 2028. Emerging hardware attacks exactly this primitive: analog in-memory computing, silicon photonics, ferroelectric devices and stochastic computing perform the underlying matrix–vector products at lower energy than digital CMOS, under conditions on scale and precision, but return a perturbed result. The usual objection is that nobody can say in advance how much perturbation a model tolerates. In EMMA, this opportunity is pursued through a PCM-based photonic matrix–vector engine as the primary demonstrator, complemented where appropriate by FeFET-based analog computing, so that device-level non-idealities are treated as explicit design parameters rather than as after-the-fact implementation losses. Diffusion models are an unusual case in which that question has a mathematical answer: their convergence theory bounds the discrepancy between the generated and target distributions by three terms, of which only one, the L 2 error of the learned score, depends on the hardware. That side is now in good shape: for stochastic samplers, bounds linear in the dimension under minimal assumptions on the data; for deterministic samplers, a matching theory under regularity. What is missing is the bridge. The theory is parameterised by a score error, the hardware literature by conductance noise, effective number of bits and photons per multiply–accumulate; nobody has written the map between them, in either direction. The quantisation literature for diffusion models reports image-quality scores; the induced L 2 score error, in the form the convergence theory needs, is directly measurable and goes unreported. One open sub-question sets the direction: the distinction between an error redrawn at each evaluation, as in stochastic computing, and an error frozen at write time, as in deterministic quantisation, has been settled in one setting only, and not the one hardware lives in; carrying it over is what this position exists to do. The work is judged on one concrete outcome. For at least one architecture, we want to show that theenergy gain from moving the score evaluation off digital CMOS is real, measured where it survives integration and not extrapolated from a single tile, and that it is paid for by a loss in sample quality that is small and, above all, tunable: a knob the designer sets, whose position the theory predicts rather than discovers afterwards.
What you will do
The work is organised in four packages, summarised below; the full research programme is available from the supervisors on request.
- Make the score error measurable, starting from Gaussian targets where the whole error decomposition is closed-form. You will report the induced score error for a given perturbation of the network’s arithmetic, the number the theory needs and the quantisation literature does not report.
- Map device physics onto the theory. Propagate the published error models for the technologies INL works on through the score network into an explicit score-error budget, separating what is proved from what is measured.
- Extend the guarantees. Existing bounds assume a deterministic oracle, or a benign randomised one; extend them to error that is part frozen at write time and part redrawn at each call, and settle, on the way, which process the score error should be controlled along.
- Co-design with INL, and deliver the proof of concept. Turn the score-error budget into an energy– accuracy frontier for INL’s PCM-based photonic matrix–vector engine, using device-calibrated circuitlevel models; a FeFET-based analog alternative may be considered where appropriate. Jointly vary sampler steps, stochastic bitstream length and effective precision to identify an operating point predicted to yield a real energy advantage over digital CMOS at a small, tunable loss in sample quality. The result is a frontier predicted from mathematical and hardware models, then verified against calibrated measurements.
A clear negative answer is more useful to us than a vague positive one: a negative result on point 3 leaves points 1 and 2 intact and turns point 4 from a prediction into a measurement.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
