Context <p>Organic solar cells (OSCs) offer a promising route toward flexible and sustainable energy technologies, yet predictive modeling of device parameters remains challenging due to the chemical diversity of donor–acceptor systems and morphology-dependent effects. In this work, we present the first systematic demonstration of using autoencoder-compressed molecular fingerprints with tree-based machine learning models to predict key OSC performance metrics—power conversion efficiency (PCE), open-circuit voltage (V<sub>oc</sub>), short-circuit current (J<sub>sc</sub>), and fill factor (FF)—from a broad experimental dataset of &#xa0;2500 donor–acceptor pairs, including both fullerene and non-fullerene acceptors. These compact models, trained on compressed descriptors of only 32 dimensions, achieved strong predictive accuracy (Pearson <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="894_2025_6514_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="61" /> </InlineMediaObject> <EquationSource Format="TEX">\(r &gt; 0.95\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>r</mi> <mo>&gt;</mo> <mn>0.95</mn> </mrow> </math></EquationSource> </InlineEquation>, <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="894_2025_6514_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="89" /> </InlineMediaObject> <EquationSource Format="TEX">\(MAE &lt; 0.4\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>M</mi> <mi>A</mi> <mi>E</mi> <mo>&lt;</mo> <mn>0.4</mn> </mrow> </math></EquationSource> </InlineEquation>, <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="894_2025_6514_Article_IEq3.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="109" /> </InlineMediaObject> <EquationSource Format="TEX">\(RMSE &lt; 0.95\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>R</mi> <mi>M</mi> <mi>S</mi> <mi>E</mi> <mo>&lt;</mo> <mn>0.95</mn> </mrow> </math></EquationSource> </InlineEquation>) while remaining lightweight enough to run on standard computing hardware. As a complementary result, some <i>k</i>-nearest neighbor models achieved near-perfect correlations (<InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="894_2025_6514_Article_IEq4.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="61" /> </InlineMediaObject> <EquationSource Format="TEX">\(r \sim 0.99\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>r</mi> <mo>∼</mo> <mn>0.99</mn> </mrow> </math></EquationSource> </InlineEquation>) and quite small errors (<InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="894_2025_6514_Article_IEq5.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="105" /> </InlineMediaObject> <EquationSource Format="TEX">\(MAE &lt; 0.044\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>M</mi> <mi>A</mi> <mi>E</mi> <mo>&lt;</mo> <mn>0.044</mn> </mrow> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="894_2025_6514_Article_IEq6.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="101" /> </InlineMediaObject> <EquationSource Format="TEX">\(RMSE&lt;0.4\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>R</mi> <mi>M</mi> <mi>S</mi> <mi>E</mi> <mo>&lt;</mo> <mn>0.4</mn> </mrow> </math></EquationSource> </InlineEquation>) in general, demonstrating the surprising strength of simple, instance-based learners when sufficient descriptive features are available. Supporting analyses reveal that fullerene datasets are more easily modeled than chemically diverse non-fullerene sets, that fingerprints encode substantial structural information, and that kernel density analyses identify critical ranges of molecular weight and energy offsets for high-efficiency devices. Collectively, this study establishes compressed fingerprint descriptors as a powerful, computationally inexpensive foundation for predictive modeling in OSCs, while also showcasing the unexpected efficacy of k-NN models trained on conventional descriptors. Together, these approaches provide a scalable path toward high-throughput prediction and guided molecular design of next-generation organic photovoltaic materials.</p> Methods <p>The dataset used in this work comprises approximately 2500 experimentally characterized donor–acceptor pairs from bulk heterojunction OSCs. These include both fullerene and non-fullerene acceptor systems. For each pair, the database provides electronic descriptors, polymerization-related metrics, and the SMILES representations of the donor and acceptor molecules. Molecular fingerprints were computed from SMILES codes using the RDKit and CDK cheminformatics toolkits. A variety of machine learning models were explored, including feedforward neural networks, autoencoders for feature compression, tree-based ensemble methods, and kernel-based regression algorithms. Hyperparameter tuning was carried out using the Optuna and BayesSearchCV libraries to ensure optimal model performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine learning-driven prediction of organic solar cell performance: a data-centric approach to molecular design

  • Victor dos Reis Rodrigues,
  • Víctor de Souza Assumção Bonfim,
  • Demétrio Antônio da Silva Filho

摘要

Context

Organic solar cells (OSCs) offer a promising route toward flexible and sustainable energy technologies, yet predictive modeling of device parameters remains challenging due to the chemical diversity of donor–acceptor systems and morphology-dependent effects. In this work, we present the first systematic demonstration of using autoencoder-compressed molecular fingerprints with tree-based machine learning models to predict key OSC performance metrics—power conversion efficiency (PCE), open-circuit voltage (Voc), short-circuit current (Jsc), and fill factor (FF)—from a broad experimental dataset of  2500 donor–acceptor pairs, including both fullerene and non-fullerene acceptors. These compact models, trained on compressed descriptors of only 32 dimensions, achieved strong predictive accuracy (Pearson \(r > 0.95\) r > 0.95 , \(MAE < 0.4\) M A E < 0.4 , \(RMSE < 0.95\) R M S E < 0.95 ) while remaining lightweight enough to run on standard computing hardware. As a complementary result, some k-nearest neighbor models achieved near-perfect correlations ( \(r \sim 0.99\) r 0.99 ) and quite small errors ( \(MAE < 0.044\) M A E < 0.044 and \(RMSE<0.4\) R M S E < 0.4 ) in general, demonstrating the surprising strength of simple, instance-based learners when sufficient descriptive features are available. Supporting analyses reveal that fullerene datasets are more easily modeled than chemically diverse non-fullerene sets, that fingerprints encode substantial structural information, and that kernel density analyses identify critical ranges of molecular weight and energy offsets for high-efficiency devices. Collectively, this study establishes compressed fingerprint descriptors as a powerful, computationally inexpensive foundation for predictive modeling in OSCs, while also showcasing the unexpected efficacy of k-NN models trained on conventional descriptors. Together, these approaches provide a scalable path toward high-throughput prediction and guided molecular design of next-generation organic photovoltaic materials.

Methods

The dataset used in this work comprises approximately 2500 experimentally characterized donor–acceptor pairs from bulk heterojunction OSCs. These include both fullerene and non-fullerene acceptor systems. For each pair, the database provides electronic descriptors, polymerization-related metrics, and the SMILES representations of the donor and acceptor molecules. Molecular fingerprints were computed from SMILES codes using the RDKit and CDK cheminformatics toolkits. A variety of machine learning models were explored, including feedforward neural networks, autoencoders for feature compression, tree-based ensemble methods, and kernel-based regression algorithms. Hyperparameter tuning was carried out using the Optuna and BayesSearchCV libraries to ensure optimal model performance.