SinoTechIntel Academic Portal
Open AccessDOI: 10.29026/oea.2026.250249Original Research

Polarization-guided diffusion prior for eyeglass reflection removal

Tsinghua University

Read Executive PreviewQuick FAQ
Polarization-guided diffusion prior for eyeglass reflection removal
Graphical Abstract / Figure
Published In
Opto-Electronic Advances (光电进展)
Published:January 15, 2026Edition:Vol. 32, Issue 1 • pp. 100-112Citation:Yating CHEN et al. (2026), Opto-Electronic Advances (光电进展)
Impact Factor3.8
Strategic Intelligence Pillar
Wide-Bandgap Semiconductors: 8-Inch SiC Wafers, GaN Power HEMT & Diamond Substrates
Explore Topic Pillar

Key Takeaways & Executive Findings

  • • • PDPrior operates with zero training data and zero ground-truth reflection-free images, eliminating the dataset dependency that cripples existing polarization-based methods under unseen lighting; this enables immediate deployment in video conferencing and facial recognition without costly paired data collection. • • The method achieves artifact-free, high-fidelity reflection removal on real-world eyeglass images captured by a division-of-focal-plane polarization camera across indoor and outdoor lighting, yielding higher face image quality assessment scores for recognition than state-of-the-art methods—directly reducing recognition errors in biometric authentication. • • By alternately updating reflection and transmission variables via gradient descent at each diffusion sampling step, PDPrior embeds physical interpretability through the forward model of reflection formation, ensuring robustness where purely data-driven approaches fail under complex lighting. • • The framework extends to window photography, showcase displays, and driver monitoring, providing a general reflection removal solution that bypasses the scalability bottleneck of paired dataset acquisition and lowers deployment barriers for consumer photography and intelligent vision systems.
Weekly Academic Intelligence

China Advanced Materials & Deep-Tech Radar

Get verified English translations, SEM micrographs & open-access PDF alerts from China's leading state key laboratories delivered to your inbox every Monday at 08:00 EST.

Institutional privacy protected100% Free Open AccessUnsubscribe anytime

Abstract

Eyeglass reflection severely degrades facial feature visibility in video conferencing and facial recognition, where the captured image is a superposition of transmission and reflection layers. Existing polarization-based reflection removal methods depend on large-scale paired polarization datasets, limiting generalization to unseen lighting conditions. This work introduces PDPrior, an untrained polarization-guided diffusion prior that requires no training data and no ground-truth reflection-free images. PDPrior leverages the generative prior of a diffusion model and incorporates polarization information as guidance to control the generation process. The reflection diffusion model uses the degree of linear polarization (DoLP) to preliminarily identify reflection regions and exploits the diffusion prior of progressively darkening facial content, enabling focus on reflective areas. The transmission coefficient computed from DoLP guides transmission image generation via the physical forward model of reflection formation. During each sampling step, reflection and transmission variables are alternately updated through gradient descent based solely on the test sample, conferring adaptability to complex lighting and diverse scenes. Real-world eyeglass reflection images were collected using a division-of-focal-plane polarization camera under various indoor and outdoor lighting environments. Experimental results demonstrate that PDPrior effectively removes eyeglass reflection, producing high-fidelity face reconstructions with no visible artifacts and achieving more robust face image quality assessment scores for recognition performance. The framework generalizes to window photography, showcase displays, and driver monitoring.

1. Introduction

Video conferencing platforms such as Zoom, Microsoft Teams, Tencent Meeting, and Feishu have experienced explosive expansion, while facial recognition-based biometric authentication has become mainstream in government and finance. In these scenarios, eyeglass reflection—arising from specular highlights, ambient scene reflections, or both—occludes critical facial features, degrading visual perception and increasing recognition errors. The captured image is physically modeled as a superposition of a transmission layer and a reflection layer, making separation a blind source separation problem under unknown lighting.

Existing polarization reflection removal approaches rely on large-scale paired polarization datasets, which are expensive to acquire and fail to generalize to unseen lighting conditions. PDPrior addresses this bottleneck by requiring no training data and no ground-truth reflection-free images. It operates solely on polarization observations collected at test time, using the degree of linear polarization (DoLP) to identify reflection regions and compute transmission coefficients. A diffusion model's generative prior is guided by polarization and a self-supervised loss based on the physical forward model, with reflection and transmission variables alternately updated via gradient descent. This untrained optimization achieves artifact-free, high-fidelity reflection removal and robust face image quality assessment scores for recognition, with extensibility to window photography, showcase displays, and driver monitoring.

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Cite This Research Paper
Yating CHEN, Liangcai CAO (2026). Polarization-guided diffusion prior for eyeglass reflection removal. Opto-Electronic Advances (光电进展). https://doi.org/10.29026/oea.2026.250249
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only: The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntelare intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is the failure mechanism of PDPrior under extreme lighting conditions, such as strong specular highlights or low-light outdoor scenes?

PDPrior relies on DoLP to preliminarily identify reflection regions and compute transmission coefficients. Under strong specular highlights, DoLP may saturate, causing inaccurate reflection masks; however, the diffusion prior of progressively darkening facial content and the alternate gradient descent updates mitigate this by focusing on reflective regions. In low-light outdoor scenes, polarization observations may have reduced signal-to-noise ratio, but the method was validated on real-world eyeglass images captured under various indoor and outdoor lighting, achieving artifact-free reconstructions. No quantitative failure threshold is reported, but robustness is demonstrated across diverse lighting.

How does PDPrior achieve cost parity against legacy paired-dataset training methods, considering it requires no training data?

Legacy methods require large-scale paired polarization datasets, incurring costs for data collection, annotation, and retraining for each new lighting condition. PDPrior eliminates these costs entirely: it requires no training data and no ground-truth reflection-free images, operating solely on polarization observations collected at test time. The computational cost shifts to per-test optimization via diffusion sampling with alternate gradient descent updates, which is a one-time inference cost. This reduces deployment overhead and enables adaptation to unseen lighting without retraining, providing a favorable cost profile for video conferencing and facial recognition applications.

What are the scalability bottlenecks when deploying PDPrior on a division-of-focal-plane polarization camera for real-time video conferencing?

The primary bottleneck is the iterative diffusion sampling process, where reflection and transmission variables are alternately updated via gradient descent at each step. This per-frame optimization may challenge real-time throughput on standard hardware. However, the method operates on test-time polarization observations without training, so scaling to multiple users requires only parallel inference. The division-of-focal-plane polarization camera provides the necessary DoLP input. No specific frame rate or latency metrics are reported, but the framework's extension to driver monitoring and showcase displays suggests practical feasibility with adequate computational resources.

How does the physical forward model of reflection formation ensure interpretability compared to black-box deep learning approaches?

The forward model explicitly represents the captured image as a superposition of transmission and reflection layers. PDPrior uses DoLP to compute the transmission coefficient and guides transmission image generation via this physical model. The self-supervised loss based on the forward model alternately updates reflection and transmission variables through gradient descent, embedding physical constraints into the diffusion process. This contrasts with black-box networks that lack explicit layer separation, providing interpretability and enabling artifact-free, high-fidelity reconstructions validated on real-world eyeglass data.

What evidence supports the claim that PDPrior achieves higher face image quality assessment scores for recognition performance than state-of-the-art methods?

Extensive experimental results on eyeglass reflection data captured in both indoor and outdoor lighting conditions demonstrate that PDPrior produces high-fidelity face reconstructions with no visible artifacts. The method achieves higher face image quality assessment scores for recognition performance compared to state-of-the-art methods. These results validate effectiveness and robustness, directly addressing the occlusion problem that increases recognition errors in facial recognition-based biometric authentication. The specific metric values and statistical significance are not detailed in the provided text, but the comparative advantage is empirically established.

Related Chinese Research & Cross-Citations

Research Citation2026
Tunable Compound Eyes with Coaxial Lens-on-Lens Ommatidia for Cooperative Bi-Focal Imaging

Tunable Compound Eyes with Coaxial Lens-on-Lens Ommatidia for Cooperative Bi-Focal Imaging

Artificial compound eyes (CEs) remain inferior to insect counterparts in ommatidial spatial arrangement, size distribution, visual field adaptability, and environmental perception. This work presents a tunable bionic CE with coaxial lens-on-lens (LoL) ommatidia, inspired by Sympetrum frequens, integrating a flexible polydimethylsiloxane (PDMS) LoL array with a microfluidic chip to achieve simultaneous bi-focal imaging. The LoL array was fabricated via femtosecond laser dual-modification of quartz glass, two-step wet etching, and soft lithography, yielding a concave template of approximately 2.6 mm². Integration with a microfluidic chamber enabled liquid-pressure modulation of CE configurations, producing a complete curved bi-focal plane that overcomes the limitations of single-focal-plane and regionalized nonuniform ommatidia CEs. Optical characterization confirmed stable focusing performance for both large and small ommatidia within their theoretical fields of view (FOVs). Cooperative bi-focal imaging was achieved by regulating FOV and relative positions of different LoL ommatidia through controlled injection of PDMS precursor. Large-FOV imaging and moving target monitoring were demonstrated, with reconstructed trajectories of triangular and dragonfly targets in 3D coordinates. The tunable CE with LoL ommatidia offers significant potential for particle image velocimetry, robotic vision, and virtual endoscopy, providing a scalable route to advanced micro-optical systems with adaptive visual field and depth perception.

Examine Full Data & PDF
Research Citation2026
AI-assisted metaphotonics: A Comprehensive Review of Artificial Intelligence-Driven Approaches for Metaphotonic Systems

AI-assisted metaphotonics: A Comprehensive Review of Artificial Intelligence-Driven Approaches for Metaphotonic Systems

The convergence of artificial intelligence (AI) and metaphotonics is creating a new paradigm for controlling light-matter interactions. The synergy of AI's ability to learn complex relationships in multidimensional data and provide ultra-fast inference with the capacity of metaphotonics to engineer optical properties not found in nature is unlocking a new era in computational design, real-time control, and fully automated optical systems. This review provides a comprehensive overview of state-of-the-art AI-driven approaches for metaphotonic systems. We focus on the solutions to real-world problems in accelerating metaphotonic simulations and inverse design, optical data characterization, and the development of fully integrated end-to-end AI-assisted metaphotonic systems. Finally, we provide our perspectives on the future research directions and emerging opportunities at the rapidly evolving intersection of metaphotonics and AI.

Examine Full Data & PDF
Research Citation2026
Polarization Unlocks Scene-Level 3D Imaging: A Commentary on Integration-Free Binocular-Polarization Fusion for Discontinuous Targets

Polarization Unlocks Scene-Level 3D Imaging: A Commentary on Integration-Free Binocular-Polarization Fusion for Discontinuous Targets

Scene-level high-precision 3D imaging remains constrained by the fundamental trade-off between imaging distance and depth accuracy. Polarization-based reconstruction offers pixel-level precision without this trade-off, yet conventional surface-normal integration fails on discontinuous targets where multiple objects are separated in space. Liu et al. (Opto-Electron Adv 9, 250267, 2026) demonstrate an integration-free approach that jointly and iteratively couples pixel-level surface normals from polarization with absolute scale information from binocular stereo vision under a unified mathematical optimization framework. This mutual-constraint formulation resolves discontinuous geometry and recovers true depth without normal integration. A scale-normalization strategy globally aligns and spatially calibrates multi-view measurements, eliminating scale drift during multi-frame point-cloud fusion. Experiments confirm scene-level, high-precision 3D reconstruction at video rates. The method extends reconstruction capability from isolated single objects to complex natural multi-object scenes, with direct relevance to autonomous driving, remote sensing, and complex scene perception. Remaining engineering bottlenecks include the fixed-focus architecture, which limits adaptation to natural scenes of varying scale and distance, and the absence of validated dynamic reconstruction for large-moving targets such as pedestrians and vehicles. The work establishes a practical pathway toward deployable scene-level passive polarization 3D imaging.

Examine Full Data & PDF
Research Citation2026

Optoelectronic Advances in the Hybrid Plasmonic Metasurface for Multi-Band and Wide-Spectrum Photodetection

Hybrid plasmonic metasurfaces have emerged as a pivotal platform for enhancing photodetection across multiple bands, yet their practical deployment is constrained by narrow operational bandwidth and high dark current. This study presents a comprehensive experimental investigation of a hybrid plasmonic metasurface photodetector that achieves a peak responsivity of 0.45 A/W at 1550 nm and a specific detectivity of 1.2 × 10^11 Jones, with a dark current density of 2.5 nA/cm² at room temperature. The device exhibits a broad spectral response from 400 nm to 1700 nm, with an external quantum efficiency exceeding 60% at 1300 nm. The metasurface, composed of gold nanodisks on a silicon-on-insulator substrate, leverages localized surface plasmon resonance to enhance light absorption and hot-carrier generation. Experimental results demonstrate a 3 dB bandwidth of 10 GHz and a rise time of 35 ps, enabling high-speed operation. The photodetector maintains stable performance over 1000 hours of continuous operation, with a degradation rate of less than 5%. These findings establish a viable route for multi-band, high-sensitivity photodetection in optical communication and imaging systems.

Examine Full Data & PDF
Research Citation2026
Millisecond-level electrically switchable metalens for adaptive rotational depth mapping and diffraction-limited imaging

Millisecond-level electrically switchable metalens for adaptive rotational depth mapping and diffraction-limited imaging

The intrinsic trade-off between depth-of-focus and lateral resolution in conventional optical systems constrains three-dimensional imaging in compact form factors. This work demonstrates an electrically tunable dual-mode metalens that integrates hydrogenated amorphous silicon (a-Si:H) meta-atoms with a liquid crystal (LC) modulator to independently manipulate left- and right-circularly polarized (LCP/RCP) light at 635 nm. Under LCP illumination, the metalens generates a rotating double-helix point spread function (PSF) encoding depth via rotation angle; under RCP illumination, it produces an extended depth-of-focus with a narrow PSF for high-resolution imaging. Propagation and geometric phases were co-optimized via rigorous coupled-wave analysis (RCWA), yielding high transmittance and precise phase control. Experimental characterization confirmed near-diffraction-limited lateral and axial resolutions. The integrated LC cell enables millisecond-scale polarization switching between depth-sensitive and high-resolution modes. Depth extraction was validated by correlating rotation angles of dual-image focal spots under mixed-polarization illumination, with axial displacements of Δz1 = 30.5 μm, Δz2 = 0 μm, and Δz3 = −48.3 μm corresponding to rotation angles β = −25.6°, 0°, and 16.9°, respectively. Depth-resolved imaging of a rubber-tree leaf, skeletal-muscle cross-section, and live planarian retrieved color-coded depth maps, demonstrating efficacy on complex biological tissues. This polarization-driven platform offers a compact solution for biomedical imaging, three-dimensional sensing, and adaptive optics.

Examine Full Data & PDF
Research Citation2026
Overcoming Challenges in InP-Based Quantum Dots: From Nucleation Mechanisms to High-Performance Quantum Dot Light-Emitting Diodes

Overcoming Challenges in InP-Based Quantum Dots: From Nucleation Mechanisms to High-Performance Quantum Dot Light-Emitting Diodes

Indium phosphide-based quantum dots (InP QDs) are positioned as the leading cadmium-free alternative for next-generation display and optoelectronic technologies, offering high photoluminescence quantum yield (PL QY), narrow emission spectra, and size-tunable wavelengths. Commercial deployment, however, remains constrained by synthetic and processing bottlenecks. State-of-the-art InP QD systems typically deliver PL QY below 90% and emission linewidths exceeding 35 nm, while device external quantum efficiency (EQE) and operational lifetime improve only incrementally. This review systematically examines the nucleation mechanisms governing InP core formation and evaluates optimization strategies for core/shell heterostructures, ligand engineering, and device architecture. A comprehensive analysis of recent breakthroughs in red, green, and blue InP-based quantum dot light-emitting diodes (QLEDs) is presented, with emphasis on charge transport modulation and suppression of charge leakage. Despite progress, a significant performance gap persists for practical display applications. Critical unresolved challenges include achieving high-performance electroluminescence from small QDs, mitigating imbalanced carrier injection that drives Auger recombination, Joule heating, and low recombination efficiency, elucidating luminescence and aging mechanisms, and improving blue-emitting device performance. The review concludes by outlining pathways to overcome these limitations, including fabrication of large-sized InP QDs with near-unity PL QY, enhancement of radiative recombination and light extraction efficiency, advanced characterization of degradation mechanisms, and performance enhancement of blue InP-based QLEDs.

Examine Full Data & PDF