Abstract Permeability across the epithelium is one of the major barriers to drug absorption and is one property subject to in silico prediction attempts. Prediction models provide a possibility to address absorption issues, early in drug discovery, for a large number of compounds. The aim of this study was to develop a general comprehensive partial least square projection to latent structures (PLS)‐model for prediction of Caco‐2 cell permeability using theoretically calculated descriptors suited for large virtual libraries. In order to deal with current issues of data quantity and quality the well‐established Caco‐2 cell model was used to generate accurate permeability data of apparent passive transport for a large set of structurally diverse compounds. PLS statistics was used to correlate calculated descriptors to log P app . This new prediction model for Caco‐2 cell permeability has incorporated many different descriptor types to deal with the multivariate nature of permeability. The model is designed to classify discovery compounds as low, medium or high permeable. A good statistical model was derived (R 2 =0.79, Q 2 =0.65, n=46) using 70 descriptors including lipophilicity, hydrogen bonding, polar surface area, size and charge descriptors and some nonlinear terms. The model has been tested and proved valid on two different external test sets (n=5 and n=125 respectively). Root mean square error of prediction (RMSEP) was 0.45 for the small external test set. The model predicted 82% of the compounds in the test sets as members to the correct class, 18% were classified wrong. No low permeable compounds were classified as high permeable and only one high permeable compound was classified as low permeable. With this model it has been shown that the in silico prediction models for Caco‐2 cell permeability has taken a step closer to meet the expectations of a high throughput filter tool applied in early phase drug discovery.