<p>Generative models cover various application areas, including image and video synthesis, natural language processing and molecular design, among many others<sup><CitationRef AdditionalCitationIDS="CR2 CR3 CR4 CR5 CR6 CR7 CR8 CR9 CR10" CitationID="CR1">1</CitationRef>–<CitationRef CitationID="CR11">11</CitationRef></sup>. As digital generative models become larger, scalable inference in a fast and energy-efficient manner becomes a challenge<sup><CitationRef AdditionalCitationIDS="CR13" CitationID="CR12">12</CitationRef>–<CitationRef CitationID="CR14">14</CitationRef></sup>. Here we present optical generative models inspired by diffusion models<sup><CitationRef CitationID="CR4">4</CitationRef></sup>, where a shallow and fast digital encoder first maps random noise into phase patterns that serve as optical generative seeds for a desired data distribution; a jointly trained free-space-based reconfigurable decoder all-optically processes these generative seeds to create images never seen before following the target data distribution. Except for the illumination power and the random seed generation through a shallow encoder, these optical generative models do not consume computing power during the synthesis of the images. We report the optical generation of monochrome and multicolour images of handwritten digits, fashion products, butterflies, human faces and artworks, following the data distributions of MNIST<sup><CitationRef CitationID="CR15">15</CitationRef></sup>, Fashion-MNIST<sup><CitationRef CitationID="CR16">16</CitationRef></sup>, Butterflies-100<sup><CitationRef CitationID="CR17">17</CitationRef></sup>, Celeb-A datasets<sup><CitationRef CitationID="CR18">18</CitationRef></sup>, and Van Gogh’s paintings and drawings<sup><CitationRef CitationID="CR19">19</CitationRef></sup>, respectively, achieving an overall performance comparable to digital neural-network-based generative models. To experimentally demonstrate optical generative models, we used visible light to generate images of handwritten digits and fashion products. In addition, we generated Van Gogh-style artworks using both monochrome and multiwavelength illumination. These optical generative models might pave the way for energy-efficient and scalable inference tasks, further exploiting the potentials of optics and photonics for artificial-intelligence-generated content.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Optical generative models

  • Shiqi Chen,
  • Yuhang Li,
  • Yuntian Wang,
  • Hanlong Chen,
  • Aydogan Ozcan

摘要

Generative models cover various application areas, including image and video synthesis, natural language processing and molecular design, among many others111. As digital generative models become larger, scalable inference in a fast and energy-efficient manner becomes a challenge1214. Here we present optical generative models inspired by diffusion models4, where a shallow and fast digital encoder first maps random noise into phase patterns that serve as optical generative seeds for a desired data distribution; a jointly trained free-space-based reconfigurable decoder all-optically processes these generative seeds to create images never seen before following the target data distribution. Except for the illumination power and the random seed generation through a shallow encoder, these optical generative models do not consume computing power during the synthesis of the images. We report the optical generation of monochrome and multicolour images of handwritten digits, fashion products, butterflies, human faces and artworks, following the data distributions of MNIST15, Fashion-MNIST16, Butterflies-10017, Celeb-A datasets18, and Van Gogh’s paintings and drawings19, respectively, achieving an overall performance comparable to digital neural-network-based generative models. To experimentally demonstrate optical generative models, we used visible light to generate images of handwritten digits and fashion products. In addition, we generated Van Gogh-style artworks using both monochrome and multiwavelength illumination. These optical generative models might pave the way for energy-efficient and scalable inference tasks, further exploiting the potentials of optics and photonics for artificial-intelligence-generated content.