Psychiatry Investig Search

CLOSE


Psychiatry Investig > Volume 23(7); 2026 > Article
Jang, Park, Shin, Lee, Ryu, Bang, Seo, Kim, Jung, Kwon, and Seok: Development of a Machine Learning-Based Model for Classifying Depression Using Physiological and Psychological Indicators

Abstract

Objective

This study aimed to develop and evaluate machine learning-based models for classifying depression symptoms using salivary hormone markers (cortisol and dehydroepiandrosterone [DHEA]) and psychological indicators from depression screening devices, and to determine the extent to which salivary hormone markers contribute to depression symptom classification.

Methods

Two models were developed, a multi-class classification (four degrees of depression) and a binary classification model (normal vs. depression), using psychological indicators that assess depression symptoms, protective and vulnerability factors, and physiological indicators including salivary cortisol and DHEA. Each model was evaluated both with and without physiological indicators. Data from 368 individuals were used for model training, with 92 individuals reserved for independent testing. Out of 36 variables, including psychological indicators from the PROtective and Vulnerable factors battEry test and physiological markers (e.g., salivary cortisol and DHEA), 24 non-demographic variables with independent contributions were selected, excluding demographic information. To ensure model robustness, overfitting risks were mitigated using k-fold cross-validation, class weight balancing, and grid search optimization. Model performance was evaluated based on accuracy, precision, recall, and receiver operating characteristic values.

Results

The multi-class classification model achieved 85.9% and 76.1% accuracy with and without physiological indicators, respectively. The binary classification model (depression presence or absence) achieved 97.8% accuracy both with and without physiological indicators.

Conclusion

The machine learning models demonstrated high accuracy in depression classification, showing a noteworthy improvement in classification performance when integrating psychological and physiological markers. Future clinical trials will be conducted to secure additional clinical data to verify the model’s reliability and clinical validity.

INTRODUCTION

Depression, a major mental health challenge, affects an estimated 300 million people worldwide with a lifetime risk of 5%-17% [1,2]. It is also anticipated to emerge as the leading cause of disability worldwide by 2030.2 The prevalence of depression in South Korea is 7.7%, with the country also having the highest suicide rate globally [3]. Thus, developing accurate predictive models for depression and suicide should be considered a public health priority to facilitate novel therapeutic approaches in South Korea.
Early detection and intervention are critical for enhancing clinical outcomes in depressive disorders. Despite this, a South Korean 2021 national-level survey revealed that only 28.2% of individuals with mental disorders had visited a medical center for treatment [4]. Among those who did not seek treatment, predominant impediments were stigma associated with mental disorders and fear of discrimination [5]. Currently, some limitations currently exist in accurately diagnosing and treating depression. First, most assessment scales rely on assessing depressive symptoms [6,7]. However, evaluating depression-related protective and vulnerability factors alongside surface depression symptoms can yield more comprehensive insights into depression’s chronicity and severity, thereby offering guidance for individualized treatment [8-11].
Despite the heterogeneity of depression, its diagnosis and treatment still hinge on depression-related symptom and sign assessment, while specific objective criteria—which could facilitate early diagnosis—have yet to be established [6,7]. Extant methods for depression and suicide prediction predominantly depend on subjective self-reported measures, potentially lacking in objectivity owing to individuals’ usual reluctance to fully disclose their thoughts [12]. The lack of objective biomarkers complicates accurate diagnosis and risks oversimplifying depression by overlooking individual symptoms and treatment response variations. Nonetheless, efforts have been made to incorporate objective indicators into depression diagnosis, with biomarkers reported to have the potential to identify treatment responsiveness and thereby improve understanding and management of depression. Meta-analytic findings indicate that cortisol is the only marker extensively studied for its ability to predict major depressive disorders’ (MDD) onset, relapse, and recurrence [13].
The pathogenesis of depression has been associated with hypothalamic-pituitary-adrenal (HPA) axis dysregulation, which leads to cortisol release as the end-product of a stress response [14,15]. Cortisol Awakening Response (CAR) includes a sharp increase in cortisol secretion after waking up, reaching its peak within 30-45 minutes and returning to baseline levels within 60 minutes after waking up [16]. Recently, CAR has been investigated as a biomarker of HPA axis function, reflecting responsiveness to external stressors upon awakening [17,18]. CAR is also useful to identify consistent stress and time conditions, rendering it useful as a diagnostic biomarker. Studies examining the CAR and depression have reported heterogeneous patterns [19,20], with depression associated with both increased and decreased CAR. These discrepancies may be attributed to differences in depression severity and differences in depressive states [21]. For example, in acute and mild-to-moderate depression groups, researchers observed an inverted U-shape with hyperactivity, whereas this inverted U-shape occurred for hypoactivity in the severe depression group [22]. A low CAR has also been associated with severe depression [23], being related to a more chronic course of the disorder [23-25].
To address these past limitations, we developed Minds.NAVI, a holistic assessment software program for depression screening that integrates psychological and biological biomarkers [26]. Particularly, it incorporates results from a psychological assessment battery, including assessments of current depressive symptoms, the PROtective and Vulnerable factors battEry (PROVE) test, and HPA axis function assessed via salivary cortisol and dehydroepiandrosterone (DHEA) levels.
Conventional depression clinical screening is insufficient to capture the holistic complexity of the disorder. Meanwhile, artificial intelligence (AI) techniques offer significant advantages through their ability to process large volumes of heterogeneous data. AI application may indeed allow a deeper understanding of depression, and assist mental health professionals in predictive decision-making [27]. AI-based predictions enable the early identification of high-risk medical conditions in patients, promoting early intervention adoption [28]. Particularly, machine learning techniques excel in feature selection and extraction across multiple data modalities (e.g., biomarkers and psychological data), which is particularly valuable when handling diverse data types [29,30]. They have demonstrated robust performance in psychiatric data analysis through advantages in pattern and relationship identification across multiple variables [29,30]. Such a systematic, multimodal approach to feature analysis might be critical for discovering meaningful patterns among various elements like salivary stress hormones, depressive symptoms, and protective-vulnerability factors. The outcome would be a more in-depth analysis of depression (e.g., discovering subtypes and allowing treatment personalization) compared with traditional assessment methods. Machine learning can also be advantageous for handling smaller datasets by leveraging cross-validation and regularization to prevent overfitting and improve model generalization [31].
In this research, we conducted a comparative analysis using four distinct models: two binary classification models (with and without salivary hormone data) to determine the presence or absence of depression according to Minds.NAVI results; two multiclass classification models (with and without physiological signals) aligned with the four-level classification of Minds.NAVI. The goal is delivering a comprehensive analysis of the impact of incorporating physiological signals alongside psychological factors in both binary and multiclass depression classification scenarios.

METHODS

Participants

Study design and data integration

This study utilized an integrated database consisting of 480 participants for the development of an AI-based mental health evaluation model. The dataset was constructed by merging data from three distinct sources:
1) Confirmatory Clinical Trial Cohort: participants recruited from multiple tertiary hospitals, including Gangnam Severance Hospital, Hallym University Sacred Heart Hospital, CHA Bundang Medical Center, and Yonsei University Wonju Severance Christian Hospital.
2) Exploratory Clinical Trial Cohort: participants recruited between February and September 2022 through Gangnam Severance Hospital and online advertisements.
3) Community-based Pilot Project Cohort: 293 participants assessed through the Minds.NAVI platform and the CHEEU Counseling Center.
The final dataset for AI training and validation comprised 167 clinical trial participants (96 healthy controls and 71 individuals with depression) and 293 community-based participants.

Eligibility and selection criteria

The selection criteria were categorized based on the nature of the recruitment source to ensure both clinical diagnostic accuracy and real-world applicability.

Clinical trial cohorts (confirmatory and exploratory)

For the clinical trial participants, eligibility was strictly screened by board-certified psychiatrists and clinical psychologists based on the following criteria:
• Depression group: individuals aged 19-50 years who met the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition diagnostic criteria for MDD and scored ≥11 on the Korean version of the Quick Inventory of Depressive Symptomatology-Self-Report (K-QIDS-SR).
• Healthy control group: individuals aged 19-50 years with no psychiatric diagnosis and a K-QIDS-SR score <5.
• Exclusion criteria: 1) physical or adrenal illnesses affecting mood (e.g., thyroid abnormalities, uncontrolled diabetes); 2) use of medications influencing depressive symptoms (e.g., hormonal agents, steroids) within the last 3 months; 3) major psychiatric comorbidities (e.g., schizophrenia, bipolar disorder); 4) suspected intellectual disability (Short Form Intelligence Test score <70); 5) pregnancy or lactation; and 6) inability to provide informed consent.

Community-based pilot project cohort

In contrast to the clinical cohorts, the pilot project adopted broader inclusion criteria to capture a diverse range of mental health states in a community setting:
• Participants: this cohort included individuals who voluntarily sought to assess their mental health or were experiencing various mental health challenges in daily life.
• Age range: the subjects in this cohort were aged between 21 and 58 years.

Ethics and consent

Written informed consent was obtained from all participants. The study protocols were approved by the Institutional Review Boards (IRB) of all participating institutions, including Gangnam Severance Hospital (No. 3-2021-0085, 3-2021-0440, 3-2021-0451), Hallym University Sacred Heart Hospital (No. 2023-10-006), CHA Bundang Medical Center (No. 2023-10-243), and Yonsei University Wonju Severance Christian Hospital (No. CR223020).

Assessment: Minds.NAVITM

A mental health assessment was carried out utilizing Minds.NAVITM, a platform developed by Minds.AI Co. Ltd. (Seoul, South Korea), which includes a self-report questionnaire for assessing depression-related protective and vulnerability factors (i.e., the PROVE test) and a salivary hormone level analysis.

PROVE

The PROVE test is a comprehensive self-report questionnaire designed for depression screening. It evaluates various depression-related aspects and incorporates protective and vulnerability factors, providing a holistic assessment to inform the overall disease course and treatment plan for individuals with depression. The validity and reliability of this questionnaire were established through comparative analyses with widely used standardized scales in the field [32].

PROVE-DS: depressive symptomatology

This subdomain includes 15 questions rated on a 5-point Likert scale (0 to 4), with higher scores indicating more severe depressive symptoms. The Cronbach’s α was 0.93 in this study.

PROVE-SR: suicide risk

This subdomain assesses suicidal ideation and risk with six questions, five being yes/no questions and one being rated on a 4-point Likert scale (1 to 4). The total score range is 0-20, with higher scores indicating higher suicide risk.

PROVE-AAT: adult attachment type

This subdomain explores adult attachment based on participants’ current close relationships using two nine-item subscales: attachment anxiety, referring to preoccupation with attachment figures and fear of rejection or abandonment; attachment avoidance, referring to discomfort with closeness. Items are rated on a 7-point Likert scale. The total score range is 9-63, with higher scores indicating greater attachment anxiety or avoidance. Adult attachment types are classified into four categories based on the subscale scores: 1) secure (low anxiety and avoidance); 2) dismissing (low anxiety, high avoidance); 3) preoccupied (high anxiety, low avoidance); and 4) disorganized (high anxiety and avoidance). The Cronbach’s α values for attachment anxiety and avoidance were 0.93 and 0.77, respectively, in this study.

PROVE-ACE: adverse childhood experience

This subdomain assesses adverse childhood experiences (ACEs). It comprises 52 items divided into six subscales: emotional abuse, physical abuse, sexual abuse, neglect, exposure to domestic violence, and bullying. Items are rated on a 4-point Likert scale, with higher scores indicating more negative childhood experiences. Participants were classified as having experienced ACE if any item exceeded the subscale cutoff. The Cronbach’s α was 0.95 for the total scale and 0.86 for emotional abuse, 0.88 for physical abuse, 0.92 for sexual abuse, 0.9 for neglect, 0.93 for domestic violence, and 0.9 for bullying.

PROVE-MC: mentalization capacity problem

This subdomain aims to identify mentalization problems through 16 items divided into five subscales: lack of emotional awareness, lack of emotional expression and interaction, psychic equivalence mode, hasty and incomplete mentalizing, and lack of mentalizing others. Items are rated on a 5-point Likert scale, with higher scores indicating greater mentalization capacity problems. The Cronbach’s α values for the subscales ranged from 0.47 to 0.76.

PROVE-KRQ: resilience

This subdomain uses 53 items to assess resilience across three subscales: self-regulation ability, interpersonal relationship ability, and psychological optimism. Items are rated on a 5-point Likert scale, with higher scores indicating greater resilience. The Cronbach’s α was 0.92 for the total scale and in the range of 0.83-0.89 for the subscales.

Salivary hormone analysis

The salivary biomarkers analyzed in this study included several indices of cortisol and DHEA to assess endocrine function. Total morning cortisol (cor_sum) was calculated to represent the cumulative cortisol secretion during the morning period. The CAR (cor_inc) was defined as the increase in cortisol levels during the first 30 minutes after waking. To evaluate the circadian rhythm of cortisol, the diurnal cortisol slope (cor_dif) was measured as the rate of decline in cortisol levels throughout the day. In addition to cortisol, total morning DHEA (dhea_sum) was assessed to measure the secretion of DHEA. Finally, the cortisol-DHEA ratio (C/D_ratio) was calculated by dividing the total morning cortisol by the total morning DHEA to determine the balance between these two hormones. Participants were instructed to collect their saliva samples at three time points: immediately upon waking (0 minutes), 30 minutes, and 60 minutes after waking, on a day with a typical level of stress. Saliva was collected without external stimulation and involved only muscle movement and expectoration into a collection tube (Simport Inc.), with a minimum volume of 2 mL at each time point. Participants were asked to refrain from smoking, eating, drinking, or brushing their teeth for 30 minutes before saliva collection. Cortisol and DHEA in saliva were analyzing according to the guidance in prior research [26,33], with some analyses being conducted using the Enzyme-Linked Immunosorbent Assay method and others using the radioimmunoassay method.

Group classification algorithm

The PROVE battery, as an integrated mental health assessment, allows for categorizing individuals into four depression levels through a two-stage evaluation process. The first stage analyzes results from PROVE-AAT, PROVE-ACE, PROVEMC, and PROVE-KRQ—all factors influencing depression development. Based on these subdomains, the balance between protective and vulnerability factors is classified as “good” (no vulnerability factors), “normal” (one vulnerable factor), or “caution” (two or more vulnerable factors). The second stage categorizes participants into depressive or non-depressive subgroups, with or without suicide risk, based on PROVE-DS and SR severity. These results are combined with salivary stress hormone analysis to produce the final Minds.NAVI results, and then categorized into four levels, as described: green (healthy), encompassing symptom-free individuals with protective factors and favorable salivary results; yellow (concern), comprising those without major depressive symptoms but with hormonal irregularities or mild symptoms; orange (mild), including individuals with mild symptoms and salivary results indicating chronic stress, or with moderate symptom severity; red (severe), consisting of those with depressive disorders (Figure 1).

Machine learning model development

Libraries and dataset

The dataset encompassed data of 480 individuals, of which 368 cases (Minds.NAVI 293+clinical 75) were used for model training and 92 clinical trial cases (46 normal, 46 depressed) were used as the test set (approximately 20%). The test set was derived from real clinical patient data, enabling direct comparisons with actual outcomes and strengthening finding reliability.
Each record contained 24 independent variables and 1 dependent variable, categorized into psychological and salivary hormone markers—cortisol and DHEA—(Table 1). These variables were collected through the Minds.NAVI test. Nineteen psychological indicators were measured by PROVE, while the salivary hormone markers related to cortisol and DHEA encompass five variables. The dependent variable is the final result of the Minds.NAVITM, which indicates depression level according to the aforementioned green to red categories.
Participants’ age ranged from 20-50 years and mostly comprised female (n=319; male, n=141), indicating a significant imbalance and the presence of missing data in the dataset. Accordingly, we excluded demographic statistics from the training data used in the classification models (Table 2).

Model and tunings

Due to the limited data size and the need for interpretability of the individual factors, we decided to evaluate prediction performance using machine learning techniques, as they have a relatively lower risk of overfitting. Among machine learning classification models, we tested the model using Random-ForestClassifier, which enables multiclass classification, along with GridSearchCV for optimal parameter selection. As the collected data show class imbalance across categories in the training set (red, 153; orange, 86; yellow, 96; green, 33) corresponding to the distribution of the 368 training samples, we employed the class_weight option, which automatically adjusts class imbalance by setting weights to appropriately reflect the importance of each class during training. This approach reduces the influence of classes with abundant data while increasing the impact of classes with limited data during the learning process. The weights were calculated as follows:
wi=nsamplesnsamples×ni
where wi is the weight of class i; nsamples is the total number of samples; nclasses is the total number of classes; ni is the number of samples in class i.
The model training set utilized k-fold cross-validation (k=5) and was trained based on accuracy metrics. Using GridSearch parameters (n_estimators, max_depth, min_samples_split, min_samples_leaf, and max_features), we generated 162 candidate models, resulting in 810 total fitting iterations. The model’s final classification output (final_val) comprises the four categories of green (healthy), yellow (concern), orange (mild depression), and red (severe depression). We evaluated both multiclass classification performance across the four categories and binary classification performance (green and yellow as normal; orange and red as abnormal). For both multiclass and binary classifications, we compared classification performance according to presence or absence of salivary test. Feature importance was calculated using the mean decrease in impurity (MDI) method, which is the default approach implemented in scikit-learn’s RandomForestClassifier. In this method, each feature’s importance is computed by averaging the reduction in Gini impurity across all decision trees in the ensemble whenever that feature is used for splitting.

Data processing

The dataset comprised 480 participants with 24 variables and had a very low percentage of missing values. No missing data were found in the clinical trial dataset. However, 68 missing values were identified in the Minds.NAVI platform data, distributed as follows: mc_ex (1), mc_ot (1), res_sc (22), res_ir (22), and res_po (22). To handle the missing values without reducing the dataset size, the “fillna” function from the pandas library was used, such that missing values were filled with the mean of each respective column. Mean imputation was chosen over more complex methods such as multiple imputation for the following reasons: 1) missing values constituted only approximately 14% of the dataset (68 out of 480 samples) and were concentrated in a single domain (resilience variables: res_ sc, res_ir, and res_po with 22 missing values each); 2) physiological markers had no missing values; and 3) given the limited sample size, more complex imputation methods could introduce additional model assumptions or artificial variance, potentially outweighing any benefit. Resilience indicators exhibit a negative correlation with depression and represent one of the key variables influencing the final classification outcome; however, applying more complex imputation methods in a limited sample environment was judged to carry the risk of introducing alternative sources of bias. We therefore chose mean imputation to maintain structural simplicity and interpretability while minimizing unnecessary assumptions. We acknowledge that mean imputation may artificially reduce variance. Sensitivity analyses examining the impact of different imputation strategies on model performance are warranted in future studies with larger datasets.
Of the total dataset, 92 participants (approximately 20% of clinical trial data; 46 normal, 46 depression; balanced distri-bution was used to ensure unbiased evaluation of final model performance) were used as the test set, while 368 participants (80%; 293 from Minds.NAVI platform data, 75 from additional clinical trials) were used for model training. The training dataset was divided using 5-fold cross-validation to ensure robust model development and prevent overfitting.

Data distribution

Regarding dataset distribution, clinical trial data included 167 participants (normal, 96; depression, 71); exploratory clinical data (FA Set) included 47 participants (normal, 12; depression, 35), confirmatory clinical data included 120 participants (normal, 84; depression, 36), and Minds.NAVI platform pilot project data included 293 participants. Regarding depression severity distribution, 14 participants were in the green, 69 participants in the yellow, 80 participants in the orange, and 130 participants in the red category. Regarding depression diagnosis, 83 participants were healthy and 210 had depressive disorders.
Regarding the training and test sets, the training set included 368 participants (Minds.NAVI, 293+additional clinical, 75), while the test set included 92 participants (approximately 20% of the total clinical trial data) from clinical trials (normal, 46; depression, 46).

RESULTS

Multiclass classification model performance

Confusion matrix

The confusion matrices and derived performance metrics showed clear performance differences between models with and without salivary data (Figure 1). The model with salivary data achieved higher overall performance, demonstrating strong sensitivities across classes (green, 0.85; yellow, 0.85; orange, 0.4; red, 1.0) together with robust specificities (green, 0.97; yellow, 0.93; orange, 0.99; red, 0.91).
The model without salivary data exhibited reduced performance (sensitivities: green, 0.81; yellow, 0.45; orange, 0.4; red, 1.0; specificities: green, 0.88; yellow, 0.93; orange, 0.96; red, 0.89). The most notable degradation occurred in the yellow category, where sensitivity dropped from 0.85 to 0.45 when salivary data were removed.
A closer examination of the misclassification patterns revealed that errors were predominantly concentrated among adjacent severity categories. In the model with salivary data, orange-class misclassifications were distributed between yellow and red categories, suggesting that the model’s difficulty lies in distinguishing clinically adjacent states rather than producing random errors. This pattern is consistent with the inherent clinical overlap between intermediate depression severity levels, where the boundaries between “mild” and “concern” or between “mild” and “severe” categories are often clinically ambiguous.

Receiver operating characteristic curve

The receiver operating characteristic (ROC) curves highlighted the performance improvement obtained by incorporating salivary data: the model with salivary data achieved higher area under the curve (AUC) values across all classes (AUCs: green, 0.98; yellow, 0.94; orange, 0.97; red, 1.00), indicating strong discrimination capability for each severity category.
The model without salivary data showed comparatively lower AUC values (green, 0.93; yellow, 0.86; orange, 0.94; red, 0.99). Both models demonstrated strong performance in identifying severe cases (red class), but salivary data inclusion improved intermediate severity level detection accuracy (yellow and orange classes), as evidenced by the larger AUC values and more favorable ROC curve profiles (Figure 2).

Feature importance

Feature importance analysis revealed consistent yet distinct patterns across both models. In both cases, depression scale total score (ds_sum) was overwhelmingly the most influential predictor, showing the single highest importance value by a considerable margin. When salivary data were included, a small group of variables—specifically suicide risk (sr_sum), total morning cortisol (cor_sum), resilience positive subscale (res_po), attachment anxiety (aat_anx), the cortisol-DHEA ratio (CD_ratio), and total DHEA (dhea_sum)—formed a second tier of predictors with clearly higher importance than the remaining variables. All other features contributed only modestly, forming a lower-importance cluster.
In the model without salivary hormone data, the hierarchy after ds_sum shifted to psychological variables alone, with suicide risk emerging as the next most influential feature, followed by resilience, attachment anxiety and avoidance, and other psychosocial indicators, which collectively showed moderate but distinguishable contributions.
Overall, despite ds_sum dominating both models, the inclusion of salivary hormone data introduced biologically meaningful predictors, most notably total cortisol and the cortisol- DHEA ratio, which rose into the mid-importance tier. Therefore, physiological information adds unique, complementary predictive value beyond standard psychological measures (Figure 3).

Binary classification model performance

Confusion matrix

In the binary classification analysis (Figure 4), both models produced identical results. Confusion matrices showed that each model achieved a sensitivity of 1.00 for detecting depression (46 true positives and no false negatives) and a specificity of 0.96 (44 true negatives and 2 false positives). Therefore, both models correctly identified all depression cases and misclassified only a small number, indicating equivalent betweenmodel performance.

ROC curve

The ROC curves for the binary classification analysis showed identical results for both models (i.e., AUC=1.00), indicating perfect discrimination performance. The AUCs for the models with and without salivary data fully overlapped, reflecting equivalent classification capability (Figure 5).

Feature importance

Feature importance analysis for the binary classification models showed a similar structure across both configurations, with the depression total score (ds_sum) being the most influential feature. When salivary data were included, the next group of higher-ranked variables were suicide risk score (sr_ sum), adult attachment anxiety (aat_anx), total morning cortisol (cor_sum), and the cortisol-DHEA ratio (C/D_ratio). Other psychological and physiological measures—such as mindfulness awareness (mc_aw), resilience (res_sc, res_ir, res_ po), cortisol change indicators (cor_dif, cor_inc), and ACE measures (ace_ea, ace_pa, ace_ne)—showed smaller contributions.
In the model without salivary hormone data, the ordering after the depression total score included suicide risk, adult attachment anxiety, resilience, and mindfulness, followed by additional attachment and childhood experience variables such as ace_ea, ace_pa, aat_avo, ace_bu, ace_ne, mc_ex, and ace_ ex. Overall, both models shared a similar pattern, with total morning cortisol and the cortisol-DHEA ratio emerging only in the model with salivary data (Figure 6).

DISCUSSION

This study developed and evaluated machine learning models for classifying depression using both physiological and psychological indicator data from the Minds.NAVI assessment system. Our findings suggest the potential value of integrating salivary hormone data with psychological assessments for depression, thus addressing a key limitation in current depression assessment practices that rely primarily on subjective self-report measures. However, given the relatively small dataset (480 participants), these results should be interpreted cautiously and require validation in larger, more diverse populations before clinical implementation.
It is important to acknowledge a structural consideration regarding the relationship between the input features and the dependent variable. The primary objective of this study was to evaluate whether a machine learning approach could achieve comparable or superior predictive performance when applied to clinical variables used by the existing rule-based Minds.NAVI system. Importantly, the input variables were selected based on their clinical relevance and their independence from the outcome variable; specifically, only variables that could be considered independent of the final classification result were included as model inputs. Furthermore, the computational operations, weighting schemes, and decision rules embedded in the existing Minds.NAVI algorithm were not included in any form during the machine learning model’s training process. The RandomForest model learned non-linear, data-driven patterns across all features simultaneously, without being constrained by the predefined thresholds or decision hierarchy used in the rule-based system. Therefore, the machine learning model does not replicate or reproduce the internal algorithm; rather, it independently derives classification patterns from the raw feature values. We acknowledge that because the dependent variable (Minds.NAVI severity classification) is ultimately derived from the same domain of clinical indicators, a structural relationship between input and output inherently exists. However, this is methodologically distinct from direct circularity, in which the output computation is simply relearned. The fact that the multiclass model did not achieve perfect accuracy—particularly for intermediate severity categories (orange class sensitivity=0.40)—provides empirical evidence that the model does not merely reproduce the rule-based system but exhibits genuine variation in predictive capacity across severity levels. Nevertheless, we recognize that the current study design does not fully resolve the circularity concern, and future studies should evaluate model performance against independent clinical diagnoses (e.g., clinician-rated severity scales such as HAM-D or MADRS) as the primary outcome variable.
The multiclass classification model showed different performance levels according to salivary hormone data (85.9% vs. 76.1% accuracy), with a 9.8 percentage point improvement when physiological markers were included. This substantial enhancement suggests that physiological markers contribute meaningful information beyond what can be captured through psychological assessments alone. The improvement was particularly pronounced in distinguishing intermediate severity levels (yellow and orange categories), where the addition of salivary hormone data led to substantial increases in both sensitivity and specificity. This differential contribution across severity levels is clinically significant, as distinguishing between mild, moderate, and severe depression has important implications for treatment planning and resource allocation. The finding aligns with previous research indicating that HPA axis dysfunction [14,15], as measured through salivary hormones, may help differentiate various depression severity levels. Particularly, it reflects the heterogeneous nature of depressive disorders.
The near-perfect binary classification performance warrants careful interpretation. It is important to note that the existing rule-based Minds.NAVI system had already demonstrated high binary classification performance in a prior prospective confirmatory clinical trial: sensitivity of 97.2%, specificity of 95.2%, and accuracy of 95.8%. Given this established baseline, it is plausible that a machine learning model trained on refined data could achieve comparable or even higher performance. Additionally, the prior clinical trial evaluated performance using 120 participants, whereas the independent test set in the current study comprised 92 participants; therefore, direct comparison between the two results should be made with caution. To minimize the risk of data leakage, 20% of the total dataset— 92 samples with clinically confirmed ground truth—was separated as an independent test set entirely excluded from the model development process and used only for final performance evaluation. The contrast between high binary classification performance and reduced multiclass accuracy suggests that the high AUC in the binary task reflects the clinical distinctiveness of the two broad categories (normal vs. depression) rather than a ceiling effect from model overfitting. Meanwhile, the reduced sensitivity for intermediate severity categories (particularly orange, sensitivity=0.40) in the multiclass model indicates that genuine prediction uncertainty remains for clinically ambiguous boundary cases, further arguing against systematic overfitting or information leakage.
In contrast to the multiclass results, both models in the binary classification achieved an identical high accuracy of 97.8% (normal vs. depression). The lack of incremental benefit from salivary data in the binary task suggests a ceiling effect, where psychological indicators alone are sufficiently robust for identifying depression presence or absence. The discrepancy between the substantial improvement in multiclass classification and the equivalent performance in binary classification therefore highlights the specific utility of physiological markers: they appear essential for capturing nuanced biological gradations and distinguishing severity levels. This pattern likely reflects the complex, non-linear relationship between HPA axis function and depression severity reported previously [21,22]. Since our stress response system can exhibit either hyperactivity or hypoactivity depending on chronicity and severity, these physiological subtleties, while critical for multiclass differentiation, may not be apparent in a simple binary diagnostic framework.
Our choice of RandomForestClassifier, a traditional machine learning approach, rather than deep learning methods, was deliberate and based on several considerations. First, our relatively small dataset size (480 participants) makes traditional machine learning methods more appropriate than deep learning approaches, as the latter typically require thousands of samples to achieve optimal performance and avoid overfitting. Second, RandomForest models offer better interpretability through feature importance analysis, which is crucial for clinical applications where understanding the variables driving the predictions is essential for building trust and clinical acceptance. Third, the ensemble nature of RandomForest, which combines multiple decision trees, provides inherent robustness and reduces variance compared to single models. Therefore, our implementation of comprehensive hyperparameter tuning via GridSearchCV, combined with a 5-fold cross-validation and class weight balancing, represents a principled approach to model development that addresses the specific challenges posed by our dataset. Deep learning methods can be explored in future studies with larger datasets, while the traditional machine learning approach adopted here aligns well with current recommendations for medical AI applications emphasizing interpretability, validation rigor, and appropriate matching of method complexity to data availability.
Feature importance analysis revealed several key insights. While psychological indicators, particularly the depression total score, remained the strongest predictors, the cortisol- DHEA ratio emerged as the third most important feature in our models. This high ranking of the cortisol-DHEA ratio corroborates recent research highlighting the importance of the HPA axis function in depression. Indeed, salivary cortisol has shown potential as a biomarker for depression, which is supported by our findings that depict its practical utility in a machine learning context.
Additionally, the cortisol-DHEA ratio’s significant contribution to model performance aligns with the growing body of evidence suggesting that HPA axis dysfunction is crucial for depression pathophysiology [14,15,34]. The integration of both psychological symptoms and physiological markers in our model contributed to model performance, which reflected the complex, multifaceted nature of depression, offering a more robust framework for depression screening and severity assessment. The improved performance for depression severity identification when including salivary hormone data suggests that physiological markers might be valuable in capturing subtle variations in depression manifestations that are hard to grasp through psychological assessments alone.
This study has several limitations. First, the relatively small dataset (480 participants) limits statistical power and increases the risk of overfitting despite our use of 5-fold cross-validation, class weight balancing, and GridSearchCV parameter optimization. The class imbalance in the dataset (particularly the small number of green category cases) suggests that performance metrics should be interpreted cautiously. The absence of an external validation cohort further limits the generalizability of our findings; future studies should incorporate independent datasets from different clinical settings to validate model performance, and the lack of extensive clinical validation data means that model performance in real-world clinical settings requires further verification. Second, the demographic composition of the sample was skewed toward younger (mean age: 35.17 years) and predominantly female participants (69.3%), which may limit the applicability of our findings to broader populations, particularly older adults or male-dominant clinical groups. Third, the performance comparison between models with and without salivary hormone data was descriptive rather than supported by formal statistical testing (e.g., paired permutation tests or bootstrapped confidence intervals). The existing Minds.NAVI system’s prospective confirmatory clinical trial was evaluated using approximately 120 clinical data samples; in a retrospective analysis context such as the present study, a substantially larger sample size would be required to achieve statistically stable significance. We estimate that the overall dataset would need to be expanded to at least approximately 1,000 samples to enable adequate train/validation/test splits and reliable statistical verification. Fourth, missing data were handled using mean imputation, which may have artificially reduced variance in the resilience-related variables (res_sc, res_ir, res_po). Although the proportion of missing data was approximately 14% (68 out of 480 samples) and confined to a single domain, more complex imputation methods in a limited sample environment could introduce additional model assumptions or alternative sources of bias. The impact of different imputation strategies on model performance should be examined in future work with sufficiently expanded datasets. Fifth, feature importance was derived using the MDI method, which can be biased toward features with high cardinality or wide value ranges. Permutation-based importance and other alternative methods should be explored in subsequent analyses to ensure the robustness of feature ranking. Sixth, the term “physiological indicators” was used broadly in the original manuscript but refers specifically to salivary cortisol and DHEA measures; we have clarified this terminology throughout the revised manuscript. Finally, although the input variables were selected based on their independence from the outcome variable and the existing algorithm’s computational rules were not included in model training, the dependent variable (Minds.NAVI classification) is ultimately derived from the same domain of clinical indicators. The reported performance metrics should therefore be interpreted with awareness of this structural relationship, and future research should adopt independent clinical outcomes, such as clinician-rated severity assessments (e.g., HAM-D, MADRS), as the primary dependent variable to more rigorously evaluate the model’s true predictive capacity.
Future research should prioritize the collection of larger, multi-site datasets to enable formal statistical comparisons between models and to establish external validation cohorts. In summary, this study has significance in confirming the potential for AI-based approaches to complement or replace existing rule-based systems, even in a data-limited environment. Future studies with larger-scale data collection and external validation are planned to further verify the generalization performance and clinical utility of the model. There are also important implications from a clinical perspective. The ability to distinguish depression severity levels with reasonable accuracy could facilitate more targeted interventions and efficient mental health resource allocation. For instance, individuals in the yellow category (attention needed) might benefit from preventive interventions or watchful waiting, while those in the red category (severe) would require immediate intensive treatment. The implementation of physiological markers can help identify individuals potentially at a higher risk of progression to more severe depression, even when their psychological symptoms appear relatively mild. The non-invasive nature of salivary hormone collection also makes this approach potentially scalable for routine clinical screening, though practical considerations such as the need for standardized collection procedures and the added cost of laboratory analysis must be carefully evaluated.

Conclusions

This study advances more objective and comprehensive methods for depression assessment, showing the effective integration of psychological and physiological markers through machine learning to improve classification accuracy, especially across severity levels. The contribution of salivary hormone data, particularly the cortisol-DHEA ratio, supports incorporating HPA axis markers into diagnostic frameworks. However, these findings must be interpreted cautiously and validation in larger cohorts is required. While performance gains are promising, their clinical utility must be weighed against cost, feasibility, and patient acceptance. Ultimately, such tools should complement, not replace clinical judgment, and future research should emphasize prospective validation, real-world implementation, and integration into clinical workflows.

Notes

Availability of Data and Material

The datasets generated or analyzed during the study are available from the corresponding author on reasonable request.

Conflicts of Interest

Jeong-Ho Seok is a professor of Yonsei University and the CEO of Minds.AI which is the owner of Minds.NAVI program. Sooah Jang, Jinsoo Park, HyunKyung Shin, and Miwoo Lee are employed by Minds.AI, Co. Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Author Contributions

Conceptualization: Sooah Jang, Jinsoo Park, Jeong-Ho Seok. Data curation: Sooah Jang, Jinsoo Park, HyunKyung Shin. Formal analysis: Jinsoo Park. Funding acquisition: Sooah Jang, Jinsoo Park, HyunKyung Shin, Miwoo Lee, Jeong-Ho Seok. Investigation: Sooah Jang, Jinsoo Park, HyunKyung Shin. Methodology: Sooah Jang, Jinsoo Park. Project administration: Sooah Jang, Jinsoo Park. Resources: Vin Ryu, Minji Bang, Juneho Seo, Tae Hui Kim, Young-Chul Jung, Manjae Kwon. Software: Jinsoo Park. Supervision: Jeong-Ho Seok. Validation: Miwoo Lee, HyunKyung Shin. Visualization: Jinsoo Park. Writing—original draft: Sooah Jang, HyunKyung Shin. Writing—review & editing: Jinsoo Park, Jeong-Ho Seok.

Funding Statement

This work was supported by the Starting Growth Technological R&D Program (Post-TIPS Program, No.20337922), funded by the Korean Ministry of SMEs and Startups in 2021.

Acknowledgments

None

Notes

Use of AI tools

The authors used ChatGPT-5.1 (OpenAI, San Francisco, CA, USA) for the purpose of language editing and improving the clarity of the manuscript. The authors reviewed and edited the output as needed and take full responsibility for the final content of the paper.

Figure 1.
Confusion matrices for multiclass classification. Confusion matrices comparing the multiclass classification models with salivary hormone data (A) and without salivary hormone data (B). Each matrix displays the four depression severity categories (green, yellow, orange, and red) on both axes, with true labels on the y-axis and predicted labels on the x-axis. Cell values indicate the number of correctly classified (diagonal) and misclassified (off-diagonal) samples. Class-wise sensitivity and specificity values are shown below each matrix. The model with salivary hormone data achieved higher overall accuracy (85.9%) compared to the model without salivary hormone data (76.1%). Sensitivity values for the model with salivary hormone data were 0.85 (green), 0.85 (yellow), 0.40 (orange), and 1.00 (red); specificity values were 0.97 (green), 0.93 (yellow), 0.99 (orange), and 0.91 (red). For the model without salivary hormone data, sensitivity values were 0.81 (green), 0.45 (yellow), 0.40 (orange), and 1.00 (red); specificity values were 0.88 (green), 0.93 (yellow), 0.96 (orange), and 0.89 (red). Test set: 92 participants.
pi-2025-0472f1.jpg
Figure 2.
ROC curves for multiclass classification. ROC curves for the multiclass classification models with salivary hormone data (A) and without salivary hormone data (B). Each panel presents class-specific ROC profiles using a one-vs-rest approach for four depression severity categories: green (dark green line), yellow (yellow line), orange (orange line), and red (red line). The dashed diagonal line represents the chance-level classifier (AUC=0.50). AUC values for the model with salivary hormone data were 0.98 (green), 0.94 (yellow), 0.97 (orange), and 1.00 (red). AUC values for the model without salivary hormone data were 0.93 (green), 0.86 (yellow), 0.94 (orange), and 0.99 (red). The inclusion of salivary hormone data improved discrimination capability across all four severity categories, with the most notable improvement observed in the yellow category (AUC improvement from 0.86 to 0.94). ROC, receiver operating characteristic; AUC, area under the curve.
pi-2025-0472f2.jpg
Figure 3.
Feature importance rankings for multiclass classification models. Horizontal bar charts showing the mean decrease in impuritybased feature importance rankings for the multiclass classification models with salivary hormone data (A) and without salivary hormone data (B). Features derived from salivary hormone measures are displayed in gold bars, while features derived from psychological assessments are displayed in green bars. In both models, the depression total score (ds_sum) was the most influential predictor. When salivary hormone data were included, total morning cortisol (cor_sum), the cortisol-DHEA ratio (CD_ratio), and total DHEA (dhea_sum) emerged as notable contributors alongside psychological variables. aat_anx, adult attachment anxiety; aat_avo, adult attachment avoidance; ace_bu, bullying; ace_ea, emotional abuse; ace_ex, exposure to domestic violence; ace_ne, neglect; ace_pa, physical abuse; CD_ratio (or C/D_ratio), cortisol-DHEA ratio; cor_dif, cortisol decrease; cor_inc, cortisol increase after waking; cor_sum, total morning cortisol; dhea_sum, total DHEA; ds_sum, depression total score; mc_aw, mentalization capacity—lack of emotional awareness; mc_ex, mentalization capacity—lack of emotional expression and interaction; mc_ha, mentalization capacity—lack of harm avoidance; mc_ot, mentalization capacity—lack of other-oriented mentalization; mc_ps, mentalization capacity—lack of psychological sophistication; ms_sum, mindfulness total score; res_ir, resilience—interpersonal relationships; res_po, resilience—positivity; res_sc, resilience—self-control; sr_sum, suicide risk total score.
pi-2025-0472f3.jpg
Figure 4.
Confusion matrices for binary classification. Confusion matrices comparing the binary classification models with salivary hormone data (A) and without salivary hormone data (B). Each matrix displays two categories (normal and depression) on both axes, with true labels on the y-axis and predicted labels on the x-axis. Cell values indicate the number of correctly classified and misclassified samples. Both models achieved identical performance: sensitivity of 1.00, specificity of 0.96, and overall accuracy of 97.8%. Of 46 normal participants, 44 were correctly classified and 2 were misclassified as depression; all 46 depression participants were correctly classified. The equivalent performance between models suggests a ceiling effect in the binary classification task, indicating that psychological indicators alone are sufficiently robust for distinguishing normal from depression categories. Test set: 92 participants (46 normal, 46 depression).
pi-2025-0472f4.jpg
Figure 5.
ROC curves for binary classification. ROC curves for the binary classification of normal versus depression groups, comparing models with salivary hormone data (A) and without salivary hormone data (B). The solid blue line represents the model’s ROC profile, and the dashed diagonal line represents the chance-level classifier (AUC=0.50). Both models achieved an AUC of 1.00, indicating perfect discrimination between normal and depression groups in the test set. The near-identical ROC profiles across both models are consistent with the ceiling effect observed in the confusion matrices (Figure 4), further suggesting that the addition of salivary hormone data provides limited incremental benefit for binary classification performance. ROC, receiver operating characteristic; AUC, area under the curve.
pi-2025-0472f5.jpg
Figure 6.
Feature importance rankings for binary classification models. Horizontal bar charts showing the mean decrease in impurity-based feature importance rankings for the binary classification models with salivary hormone data (A) and without salivary hormone data (B). Features derived from salivary hormone measures are displayed in gold bars, while features derived from psychological assessments are displayed in green bars. In both models, the depression total score (ds_sum) was the most influential predictor, followed by suicide risk total score (sr_sum) and adult attachment anxiety (aat_anx). When salivary hormone data were included, total morning cortisol (cor_sum) and the cortisol-DHEA ratio (CD_ratio) formed part of the higher-ranked feature group. aat_anx, adult attachment anxiety; aat_avo, adult attachment avoidance; ace_bu, bullying; ace_ea, emotional abuse; ace_ex, exposure to domestic violence; ace_ne, neglect; ace_pa, physical abuse; CD_ratio (or C/D_ratio), cortisol-DHEA ratio; cor_dif, cortisol decrease; cor_inc, cortisol increase after waking; cor_sum, total morning cortisol; dhea_sum, total DHEA; ds_sum, depression total score; mc_aw, mentalization capacity—lack of emotional awareness; mc_ex, mentalization capacity—lack of emotional expression and interaction; mc_ha, mentalization capacity—lack of harm avoidance; ms_sum, mindfulness total score; res_ir, resilience—interpersonal relationships; res_po, resilience—positivity; res_sc, resilience—self-control; sr_sum, suicide risk total score.
pi-2025-0472f6.jpg
Table 1.
Data variables used in model training
Clinical variables Description
ds_sum Depression total score (0-48) by PROVE-DS
aat_avo Adult attachment - avoidance (0-63) by PROVE-AAT
aat_anx Adult attachment - anxiety (0-63) by PROVE-AAT
ace_ea ACEs - emotional abuse (0-3) by PROVE-ACE
ace_pa ACEs - physical abuse (0-3) by PROVE-ACE
ace_sa ACEs - sexual abuse (0-3) by PROVE-ACE
ace_ne ACEs - neglect (0-3) by PROVE-ACE
ace_ex ACEs - exposure to domestic violence (0-3) by PROVE-ACE
ace_bu ACEs - bullying (0-3) by PROVE-ACE
sr_sum SR total score (0-20) by PROVE-SR
mc_aw MC - lack of emotional awareness (0-4) by PROVE-MC
mc_ex MC - lack of emotional expression and interaction (0-4) by PROVE-MC
mc_ps MC - psychic equivalence mode (0-4) by PROVE-MC
mc_ha MC - hasty incomplete mentalizing(0-4) by PROVE-MC
mc_ot MC - lack of mentalizing others (0-4) by PROVE-MC
ms_sum MC summary score (0-28) by PROVE-MC
res_sc Resilience - self-control (low, medium, and high) by PROVE-KRQ
res_ir Resilience - interpersonal relationships (low, medium, and high) by PROVE-KRQ
res_po Resilience - positivity (low, medium, and high) by PROVE-KRQ
cor_sum Total morning cortisol
cor_inc Cortisol increase during 30 minutes after waking
cor_dif Diurnal cortisol slope
dhea_sum Total morning DHEA
C/D_ratio Cortisol-DHEA ratio (cor_sum / dhea_sum)

AAT, adult attachment type; ACE, adverse childhood experience; DHEA, dehydroepiandrosterone; DS, depressive symptomatology; KRQ, Korean resilience questionnaire; MC, mentalization capacity; PROVE, PROtective and Vulnerable factors battEry; SR, suicide risk.

Table 2.
Demographic data (N=460)
Variable Value
Age (yr) 35.17
Sex
 Male 141 (30.7)
 Female 319 (69.3)
Education
 Middle school 1 (0.2)
 High school 87 (19.0)
 Community college 21 (4.6)
 Undergraduate 191 (41.5)
 Graduate 65 (14.1)
 Not available 95 (20.6)
Marriage
 Unmarried 236 (51.3)
 Married 114 (24.8)
 Cohabit 7 (1.5)
 Divorced 7 (1.5)
 Separated 1 (0.2)
 Not available 95 (20.7)
Occupation
 Employed 160 (34.8)
 Unemployed 53 (11.5)
 Student 60 (13.0)
 Housewife 18 (3.9)
 Others 21 (4.6)
 Not available 148 (32.2)

Data are presented as mean or number (%).

REFERENCES

1. Charlson FJ, Ferrari AJ, Flaxman AD, Whiteford HA. The epidemiological modelling of dysthymia: application for the Global Burden of Disease Study 2010. J Affect Disord 2013;151:111-120.
crossref pmid
2. World Health Organization. Depression [Internet] Available at: https://www.who.int/health-topics/depression. Accessed September 11, 2025.

3. Rim SJ, Hahm BJ, Seong SJ, Park JE, Chang SM, Kim BS, et al. Prevalence of mental disorders and associated factors in Korean adults: national mental health survey of Korea 2021. Psychiatry Investig 2023;20:262-272.
crossref pmid pmc pdf
4. Ministry of Health and Welfare, National Center for Mental Health. 2021 National Mental Health Survey. Seoul: National Center for Mental Health; 2021.

5. Korea Disease Control and Prevention Agency. Korea health statistics 2020: Korea National Health and Nutrition Examination Survey (KNHANES VIII-2). Cheongju: Korea Disease Control and Prevention Agency; 2022.

6. Fried EI, Epskamp S, Nesse RM, Tuerlinckx F, Borsboom D. What are ‘good’ depression symptoms? Comparing the centrality of DSM and non-DSM symptoms of depression in a network analysis. J Affect Disord 2016;189:314-320.
crossref pmid
7. Pettersson A, Boström KB, Gustavsson P, Ekselius L. Which instruments to support diagnosis of depression have sufficient accuracy? A systematic review. Nord J Psychiatry 2015;69:497-508.
crossref pmid
8. Giano Z, Ernst CW, Snider K, Davis A, O’Neil AM, Hubach RD. ACE domains and depression: investigating which specific domains are associated with depression in adulthood. Child Abuse Negl 2021;122:105335
crossref pmid
9. Marganska A, Gallagher M, Miranda R. Adult attachment, emotion dysregulation, and symptoms of depression and generalized anxiety disorder. Am J Orthopsychiatry 2013;83:131-141.
crossref pmid
10. Fischer-Kern M, Tmej A. Mentalization and depression: theoretical concepts, treatment approaches and empirical studies - an overview. Z Psychosom Med Psychother 2019;65:162-177.
crossref pmid pdf
11. Waugh CE, Koster EH. A resilience framework for promoting stable remission from depression. Clin Psychol Rev 2015;41:49-60.
crossref pmid
12. Busch KA, Fawcett J, Jacobs DG. Clinical correlates of inpatient suicide. J Clin Psychiatry 2003;64:14-19.
crossref
13. Kennis M, Gerritsen L, van Dalen M, Williams A, Cuijpers P, Bockting C. Prospective biomarkers of major depressive disorder: a systematic review and meta-analysis. Mol Psychiatry 2020;25:321-338.
crossref pmid pdf
14. Eisenlohr-Moul TA, Miller AB, Giletta M, Hastings PD, Rudolph KD, Nock MK, et al. HPA axis response and psychosocial stress as interactive predictors of suicidal ideation and behavior in adolescent females: a multilevel diathesis-stress framework. Neuropsychopharmacology 2018;43:2564-2571.
crossref pmid pmc pdf
15. Ahn RS, Lee YJ, Choi JY, Kwon HB, Chun SI. Salivary cortisol and DHEA levels in the Korean population: age-related differences, diurnal rhythm, and correlations with serum levels. Yonsei Med J 2007;48:379-388.
crossref pmid pmc
16. Pruessner JC, Wolf OT, Hellhammer DH, Buske-Kirschbaum A, von Auer K, Jobst S, et al. Free cortisol levels after awakening: a reliable biological marker for the assessment of adrenocortical activity. Life Sci 1997;61:2539-2549.
crossref pmid
17. Wilhelm I, Born J, Kudielka BM, Schlotz W, Wüst S. Is the cortisol awakening rise a response to awakening? Psychoneuroendocrinology 2007;32:358-366.
crossref pmid
18. Wüst S, Federenko I, Hellhammer DH, Kirschbaum C. Genetic factors, perceived chronic stress, and the free cortisol response to awakening. Psychoneuroendocrinology 2000;25:707-720.
crossref pmid
19. Varela YM, de Almeida RN, Galvão ACM, de Sousa Jr GM, de Lima ACL, da Silva NG, et al. Psychophysiological responses to group cognitive-behavioral therapy in depressive patients. Curr Psychol 2023;42:592-601.
crossref pdf
20. Galvão ACM, de Almeida RN, Silva EADS, Freire FAM, Palhano-Fontes F, Onias H, et al. Cortisol modulation by ayahuasca in patients with treatment resistant depression and healthy controls. Front Psychiatry 2018;9:185
pmid pmc
21. Dedovic K, Ngiam J. The cortisol awakening response and major depression: examining the evidence. Neuropsychiatr Dis Treat 2015;11:1181-1189.
crossref pmid pmc
22. Stetler C, Miller GE. Blunted cortisol response to awakening in mild to moderate depression: regulatory influences of sleep patterns and social contacts. J Abnorm Psychol 2005;114:697-705.
crossref pmid
23. Vreeburg SA, Hoogendijk WJ, DeRijk RH, van Dyck R, Smit JH, Zitman FG, et al. Salivary cortisol levels and the 2-year course of depressive and anxiety disorders. Psychoneuroendocrinology 2013;38:1494-1502.
crossref pmid
24. Hsiao FH, Yang TT, Ho RT, Jow GM, Ng SM, Chan CL, et al. The selfperceived symptom distress and health-related conditions associated with morning to evening diurnal cortisol patterns in outpatients with major depressive disorder. Psychoneuroendocrinology 2010;35:503-515.
crossref pmid
25. Veen G, van Vliet IM, DeRijk RH, Giltay EJ, van Pelt J, Zitman FG. Basal cortisol levels in relation to dimensions and DSM-IV categories of depression and anxiety. Psychiatry Res 2011;185:121-128.
crossref pmid
26. Jang S, Kim IY, Choi SW, Lee A, Lee JY, Shin H, et al. Exploratory clinical trial of a depression diagnostic software that integrates stress biomarkers and composite psychometrics. Psychiatry Investig 2024;21:230-241.
crossref pmid pmc pdf
27. Dwyer DB, Falkai P, Koutsouleris N. Machine learning approaches for clinical psychology and psychiatry. Annu Rev Clin Psychol 2018;14:91-118.
crossref pmid
28. Sidey-Gibbons JAM, Sidey-Gibbons CJ. Machine learning in medicine: a practical introduction. BMC Med Res Methodol 2019;19:64
crossref pmid pmc pdf
29. Madububambachu U, Ukpebor A, Ihezue U. Machine learning techniques to predict mental health diagnoses: a systematic literature review. Clin Pract Epidemiol Ment Health 2024;20:e17450179315688
crossref pmid pmc pdf
30. Lucasius C, Ali M, Patel T, Kundur D, Szatmari P, Strauss J, et al. A procedural overview of why, when and how to use machine learning for psychiatry. Nat Mental Health 2025;3:8-18.
crossref pdf
31. Ghojogh B, Crowley M. The theory behind overfitting, cross validation, regularization, bagging, and boosting: tutorial. arXiv:1905.12787. 2019.

32. Lee JY, Choi SW, Jang SA, Ryu JS, Shin HK, Sim JY. [Development of the battery test for screening of depression and mental health: PROtective and Vulnerable factors battEry Test (PROVE)]. J Korean Neuropsychiatr Assoc 2021;60:143-157. Korean.
crossref pdf
33. Jang S, Choi SW, Ahn R, Lee JY, Kim J, Seok JH. Relationship of resilience factors with biopsychosocial markers using a comprehensive home evaluation kit for depression and suicide risk: a real-world data analysis. Front Psychiatry 2022;13:847498
crossref pmid pmc
34. Varghese FP, Brown ES. The hypothalamic-pituitary-adrenal axis in major depressive disorder: a brief primer for primary care physicians. Prim Care Companion J Clin Psychiatry 2001;3:151-155.
pmid pmc
TOOLS
Share:
Facebook Twitter Linked In Google+
METRICS Graph View
  • 0 Crossref
  •   Scopus
  • 865 View
  • 54 Download


ABOUT
AUTHOR INFORMATION
ARTICLE CATEGORY

Browse all articles >

BROWSE ARTICLES
Editorial Office
#522, G-five Central Plaza, 27 Seochojungang-ro 24-gil, Seocho-gu, Seoul 06601, Korea
Tel: +82-2-537-6171  Fax: +82-2-537-6174    E-mail: psychiatryinvest@gmail.com                

Copyright © 2026 by Korean Neuropsychiatric Association.

Developed in M2PI

Close layer
prev next