De-occluding broadband metalens

1,2,3,7,8 POSTECH 4,6 University of Washington 5 UNIST 9 POSCO-POSTECH-RIST Convergence Research Center 10 National Institute of Nanomaterials Technology
Nature Communications · 2026

* Equal contribution · Corresponding authors

Full affiliations
  1. Department of Computer Science and Engineering, POSTECH
  2. Department of Mechanical Engineering, POSTECH
  3. Graduate School of Artificial Intelligence, POSTECH
  4. Department of Electrical and Computer Engineering, University of Washington
  5. Department of Electrical Engineering, UNIST
  6. Department of Physics, University of Washington
  7. Department of Chemical Engineering, POSTECH
  8. Department of Electrical Engineering, POSTECH
  9. POSCO-POSTECH-RIST Convergence Research Center for Flat Optics and Metaphotonics
  10. National Institute of Nanomaterials Technology
Website teaser showing an obstructed capture, the split-spectrum metalens, the compact prototype, and the recovered scene

The learned metalens keeps far-scene light in the filter pass bands while shifting near-depth obstruction light into stop bands, preserving real scene evidence in a compact camera.

Abstract

Obstructions such as raindrops, fences, or dust degrade images, especially when cleaning is infeasible. Conventional approaches use bulky compound-lens arrays or computational inpainting, compromising compactness or fidelity. Metalenses promise compact imaging, but achieving both broadband and obstruction-free imaging remains challenging because a metalens cannot simultaneously focus distant scenes and defocus nearby occlusions across a broadband spectrum.

We introduce a de-occluding broadband metalens that suppresses obstructions while enabling broadband imaging with an optimized metalens design map. A multi-band spectral filter divides each color channel into pass and stop bands: far-scene light is focused through the pass bands, while focused light from nearby obstructions shifts into the stop bands and is rejected. A neural network further enhances the captured image.

Key Results

Source-verified optical captures show how the proposed system suppresses nearby obstructions and preserves evidence for downstream vision.

Real experiment · dirt obstruction

Recover the scene without hallucinating missing content

The obstruction itself is shown first. The hyperbolic lens is designed at 532 nm, while the broadband baseline retains near-depth dirt. Our result approaches the unobstructed reference.

Dirt obstruction used in the real experiment
Obstructionnear-depth dirt
Hyperbolic metalens result under dirt obstruction, from a lens designed at 532 nanometers
Hyperbolicdesigned at 532 nm
Broadband metalens result with visible dirt obstruction
Broadbandnon-de-occluding
Proposed de-occluding broadband metalens result with the dirt obstruction suppressed
Oursde-occluding
Unobstructed ground-truth scene for the dirt experiment
Ground truthcompound-lens reference
Vision task · object detection

VisDrone

Each optical input is paired with its detector prediction. The digital reference is paired with the dataset ground-truth annotation.

Digital GT

Original digital VisDrone reference image
Input
Dataset ground-truth annotations for VisDrone
Ground-truth label

Hyperbolic

VisDrone input captured through the hyperbolic metalens
Input
Object-detection prediction from the hyperbolic metalens capture
Prediction

Broadband

VisDrone input captured through the broadband metalens
Input
Object-detection prediction from the broadband metalens capture
Prediction

Ours

VisDrone input captured through the proposed metalens
Input
Object-detection prediction from the proposed metalens capture
Prediction
Vision task · medical segmentation

Kvasir-SEG

Polyp-segmentation predictions are compared under identical scene content and obstruction conditions.

Digital GT

Original digital Kvasir-SEG reference image
Input
Ground-truth Kvasir-SEG mask
Ground-truth label

Hyperbolic

Kvasir-SEG input captured through the hyperbolic metalens
Input
Polyp-segmentation prediction from the hyperbolic metalens capture
Prediction

Broadband

Kvasir-SEG input captured through the broadband metalens
Input
Polyp-segmentation prediction from the broadband metalens capture
Prediction

Ours

Kvasir-SEG input captured through the proposed metalens
Input
Polyp-segmentation prediction from the proposed metalens capture
Prediction
Vision task · semantic segmentation

Cityscapes

Semantic labels remain spatially coherent when the proposed optics suppress the near-depth obstruction before inference.

Digital GT

Original digital Cityscapes reference image
Input
Ground-truth Cityscapes semantic label
Ground-truth label

Hyperbolic

Cityscapes input captured through the hyperbolic metalens
Input
Semantic-segmentation prediction from the hyperbolic metalens capture
Prediction

Broadband

Cityscapes input captured through the broadband metalens
Input
Semantic-segmentation prediction from the broadband metalens capture
Prediction

Ours

Cityscapes input captured through the proposed metalens
Input
Semantic-segmentation prediction from the proposed metalens capture
Prediction

All panels are source-verified crops from Figures 5 and 6; no generated or inpainted content is used.

Videos

Because our method operates in a single shot, it naturally extends to video imaging.

Dynamic obstruction
Cube scene
Flower scene
Frog scene
Car-mounted compact cameraOutdoor input video and semantic predictions captured from a moving vehicle.

Depth–Wavelength Symmetry

For a diffractive lens, a coupled change in source wavelength and depth can produce nearly invariant PSFs.

A wavelength longer than the design wavelength causes a focal-front shift in a diffractive lens On small screens, scroll horizontally to inspect the complete focal-intensity map.
01 / 03

A longer wavelength shifts the focus forward

For λ > λd, chromatic phase mismatch produces a focal-front shift relative to the design wavelength.

From depth–wavelength symmetry to split-spectrum filtering

Step 1: monochromatic far-depth PSFs accumulate from lambda 1 through lambda 9, then combine into an x-lambda PSF scan On small screens, scroll horizontally to inspect the complete recap figure.
01 / 04

Construct the far-depth PSF wavelength scan

Stack the monochromatic far-depth PSFs across wavelength to form the x–λ scan.

Optimizing the metalens design map

Differentiable optimization pipeline for the metalens design map theta of x and y using PSF and image simulators
Metalens optimization pipeline. Only the meta-atom orientation map θ(x, y), the metalens design map, is optimized through differentiable PSF and image simulators.
Simulated sensor images at steps 50, 200, and 3200 during optimization of the metalens design map
Optimization convergence. As θ(x, y) converges, the simulated capture preserves a sharp far-depth scene while blurring near-depth obstructions.
Simulated x-lambda scans of far- and near-depth PSFs from the learned metalens design, with pass bands highlighted
Learned far- and near-depth PSF scans. In the pass bands, the learned design produces sharp far-depth PSFs and blurry near-depth PSFs; the behavior reverses in the stop bands.

Fabrication and Characterization

The visible-light SiNx metalens was fabricated, and its far- and near-depth PSFs were measured from 430 to 645 nm in 5 nm steps.

Top-down and tilted scanning electron micrographs followed by photographs of the de-occluding broadband, broadband baseline, and hyperbolic-phase metalenses
Measured normalized x-lambda PSF scans at far and near depths for the three fabricated metalenses
Measured PSF x–λ scans, normalized to total intensity. Pass bands are highlighted; focused far-depth light is transmitted, while focused near-depth light is filtered out.
Measured two-dimensional PSFs from 445 to 640 nanometers for the three metalenses at far and near depths
Measured 2D PSFs across wavelength and depth, normalized to peak intensity. The stop-band PSFs of Ours are marked with black boxes.
4 mm focal length2.516 mm apertureB 457 / G 530 / R 628 nm pass-band centersB 22 / G 20 / R 28 nm pass-band bandwidths

Quantitative Results

Image reconstruction

Static obstruction

Reconstruction PSNR

Scale: 0-25 dB
Hyperbolic15.83
Broadband18.79
Ours20.94
21-frame sequence

Dynamic reconstruction

Mean over frames
MethodPSNR ↑SSIM ↑LPIPS ↓
Hyperbolic20.720.66400.5402
Broadband22.160.69380.5203
Ours25.390.80160.4052

24.15 FPS network inference at FP16, batch 1, 512 × 512 on an RTX 6000 Ada.

Vision-task robustness

UnobstructedObstructed
Object detection

VisDrone mAP50

Scale: 0-25
Hyperbolic21.75 / 3.50
Broadband10.66 / 2.92
Ours24.27 / 17.04
Medical segmentation

Kvasir-SEG IoU

Scale: 0-100
Hyperbolic80.80 / 34.72
Broadband81.99 / 59.50
Ours83.56 / 83.17
Semantic segmentation

Cityscapes mIoU

Scale: 0-100
Hyperbolic60.01 / 46.66
Broadband59.42 / 46.01
Ours70.80 / 67.01

Each pair reports unobstructed / obstructed performance. Digital GT provides the reference image and label, not a comparable optical-method score.

Additional Results

Printed DIV2K target · dirt obstruction

Mandala

Compound-lens GT

Unobstructed compound-lens reference of the printed mandala target
Unobstructed reference

Hyperbolic

Mandala reconstruction from the hyperbolic metalens under dirt obstruction
Reconstruction

Broadband

Mandala reconstruction from the broadband metalens under dirt obstruction
Reconstruction

Ours

Mandala reconstruction from the proposed metalens under dirt obstruction
Reconstruction
Printed DIV2K target · fence obstruction

Butterfly

Compound-lens GT

Unobstructed compound-lens reference of the printed butterfly target
Unobstructed reference

Hyperbolic

Butterfly reconstruction from the hyperbolic metalens under fence obstruction
Reconstruction

Broadband

Butterfly reconstruction from the broadband metalens under fence obstruction
Reconstruction

Ours

Butterfly reconstruction from the proposed metalens under fence obstruction
Reconstruction
Dynamic obstruction · three consecutive frames

Dynamic single-shot capture

Compound-lens GT

Unobstructed compound-lens reference for the dynamic obstruction experiment
Unobstructed reference

Hyperbolic

Hyperbolic metalens dynamic reconstruction at frame one
Frame 1
Hyperbolic metalens dynamic reconstruction at frame two
Frame 2
Hyperbolic metalens dynamic reconstruction at frame three
Frame 3

Broadband

Broadband metalens dynamic reconstruction at frame one
Frame 1
Broadband metalens dynamic reconstruction at frame two
Frame 2
Broadband metalens dynamic reconstruction at frame three
Frame 3

Ours

Proposed metalens dynamic reconstruction at frame one
Frame 1
Proposed metalens dynamic reconstruction at frame two
Frame 2
Proposed metalens dynamic reconstruction at frame three
Frame 3
Controlled vision benchmark · object detection

Additional VisDrone result

Aggregate benchmark · 10 sampled validation images · Supplementary Table S8 · mAP50 · Scale 0–25

UNOBSTRUCTEDtoOBSTRUCTED

VisDrone aggregate mAP50 before and after adding the near-depth fence obstruction
MethodUnobstructed to obstructedChange
Hyperbolic
21.75% → 3.50%
−18.25 pp
Broadband
10.66% → 2.92%
−7.74 pp
Ours
24.27% → 17.04%
−7.23 pp

Digital GT

Digital VisDrone reference image for the obstructed comparison
Input
Dataset ground-truth object labels for the VisDrone comparison
Ground-truth label

Hyperbolic

Fence-obstructed VisDrone input captured through the hyperbolic metalens
Input
VisDrone prediction from the obstructed hyperbolic capture
Prediction

Broadband

Fence-obstructed VisDrone input captured through the broadband metalens
Input
VisDrone prediction from the obstructed broadband capture
Prediction

Ours

Fence-obstructed VisDrone input captured through the proposed metalens
Input
VisDrone prediction from the obstructed proposed-metalens capture
Prediction
Controlled vision benchmark · medical segmentation

Additional Kvasir-SEG result

Aggregate benchmark · 14 sampled validation images · Supplementary Table S9 · IoU · Scale 0–100

UNOBSTRUCTEDtoOBSTRUCTED

Kvasir-SEG aggregate IoU before and after adding the near-depth blood-drop obstruction
MethodUnobstructed to obstructedChange
Hyperbolic
80.80% → 34.72%
−46.08 pp
Broadband
81.99% → 59.50%
−22.49 pp
Ours
83.56% → 83.17%
−0.39 pp

Digital GT

Digital Kvasir-SEG reference image for the obstructed comparison
Input
Dataset ground-truth mask for the Kvasir-SEG comparison
Ground-truth label

Hyperbolic

Blood-drop-obstructed Kvasir-SEG input captured through the hyperbolic metalens
Input
Kvasir-SEG prediction from the obstructed hyperbolic capture
Prediction

Broadband

Blood-drop-obstructed Kvasir-SEG input captured through the broadband metalens
Input
Kvasir-SEG prediction from the obstructed broadband capture
Prediction

Ours

Blood-drop-obstructed Kvasir-SEG input captured through the proposed metalens
Input
Kvasir-SEG prediction from the obstructed proposed-metalens capture
Prediction
Controlled vision benchmark · semantic segmentation

Additional Cityscapes result

Aggregate benchmark · 10 sampled validation images · Supplementary Table S10 · mIoU · Scale 0–100

UNOBSTRUCTEDtoOBSTRUCTED

Cityscapes aggregate mIoU before and after adding the near-depth dirt obstruction
MethodUnobstructed to obstructedChange
Hyperbolic
60.01% → 46.66%
−13.35 pp
Broadband
59.42% → 46.01%
−13.41 pp
Ours
70.80% → 67.01%
−3.79 pp

Digital GT

Digital Cityscapes reference image for the obstructed comparison
Input
Dataset ground-truth semantic label for the Cityscapes comparison
Ground-truth label

Hyperbolic

Dirt-obstructed Cityscapes input captured through the hyperbolic metalens
Input
Cityscapes prediction from the obstructed hyperbolic capture
Prediction

Broadband

Dirt-obstructed Cityscapes input captured through the broadband metalens
Input
Cityscapes prediction from the obstructed broadband capture
Prediction

Ours

Dirt-obstructed Cityscapes input captured through the proposed metalens
Input
Cityscapes prediction from the obstructed proposed-metalens capture
Prediction
RoadSidewalkBuildingWallPoleTraffic lightTraffic signVegetationTerrainSkyPersonRiderCarBusMotorcycleBicycle
Compact prototype · outdoor semantic segmentation

Outdoor compact-prototype vision

Frame 1

InputOutdoor compact-camera input frame one with a near-depth obstruction
PredictionSemantic prediction for outdoor compact-camera frame one with obstruction

Frame 2

InputOutdoor compact-camera input frame two with a near-depth obstruction
PredictionSemantic prediction for outdoor compact-camera frame two with obstruction

Frame 3

InputOutdoor compact-camera input frame three with a near-depth obstruction
PredictionSemantic prediction for outdoor compact-camera frame three with obstruction

BibTeX

@article{yoon2026deoccluding,
  title   = {De-occluding broadband metalens},
  author  = {Yoon, Seungwoo and Kang, Dohyun and Choi, Eunsue and Lee, Sohyun and Kim, Seoyeon and Choi, Minho and Heo, Hyeonsu and Shin, Dong-Ha and Kwak, Suha and Majumdar, Arka and Rho, Junsuk and Baek, Seung-Hwan},
  journal = {Nature Communications},
  year    = {2026},
  doi     = {10.1038/s41467-026-76606-0},
  url     = {https://doi.org/10.1038/s41467-026-76606-0}
}
Expanded project figure