Doctor of Technical Sciences, Associate Professor
International Islamic Academy of Uzbekistan
Uzbekistan, Tashkent
Email: raxmanov@gmail.com
FORMATION OF CLINICAL-LABORATORY INDICATORS AND A DATABASE FOR THE DETECTION OF HEPATITIS B AND C: A DIGITAL APPROACH IN THE CONTEXT OF UZBEKISTAN
УДК : 616.36-002:004.8
Abstract
One of the most common liver diseases in adults and children worldwide is viral hepatitis. Hepatitis B and C in particular raise the risk of cirrhosis, fibrosis, and hepatocellular carcinoma, among other long-term complications. Preventing chronic liver damage, especially in children and adolescents, requires early detection of these diseases. The systematization of the clinical and laboratory indicators required for the detection of hepatitis B and hepatitis C, the creation of a database structure based on these indicators, and the design of a digital software module concept relevant to Uzbekistan are all covered in this article. The study is based on the findings of a nationwide screening program that was carried out in 14 Uzbek regions between July 2022 and June 2024. In this screening program, 1,048,575 individuals were examined, and positive results were recorded in 2.89% for HBsAg and 3.52% for anti-HCV [1]. On this basis, the article proposes a conceptual design for a unified electronic database that integrates the patient’s demographic, clinical, laboratory, virological, and instrumental indicators, automatically computes derived non-invasive indices (AST/ALT ratio, APRI, and FIB-4), and stratifies patients into risk groups using transparent, threshold-based (rule-based) criteria drawn from established clinical guidelines rather than a trained machine-learning model. The proposed framework is intended to serve as a foundation for the early detection of hepatitis B and C, the standardization of monitoring, and, as a future direction, the development of validated machine-learning-based clinical decision-support systems.
Аннотация
Одним из наиболее распространённых заболеваний печени у взрослых и детей во всём мире является вирусный гепатит. В частности, гепатиты B и C повышают риск развития цирроза, фиброза и гепатоцеллюлярной карциномы, а также других долгосрочных осложнений. Для предотвращения хронического поражения печени, особенно у детей и подростков, необходимо раннее выявление этих заболеваний. В данной статье рассматриваются систематизация клинико-лабораторных показателей, необходимых для выявления гепатитов B и C, формирование структуры базы данных на их основе, а также разработка концепции цифрового программного модуля, актуального для условий Узбекистана. Исследование основано на результатах общенациональной программы скрининга, проведённой в 14 регионах Узбекистана в период с июля 2022 года по июнь 2024 года. В ходе данного скрининга было обследовано 1 048 575 человек, при этом положительные результаты были зарегистрированы у 2,89% по HBsAg и у 3,52% по anti-HCV [1]. На этой основе в статье предлагается концептуальная структура единой электронной базы данных, объединяющей демографические, клинические, лабораторные, вирусологические и инструментальные показатели пациента, автоматически рассчитывающей производные неинвазивные индексы (соотношение АСТ/АЛТ, APRI и FIB-4) и распределяющей пациентов по группам риска на основе прозрачных пороговых (rule-based) критериев, опирающихся на действующие клинические рекомендации, а не на обученную модель машинного обучения. Предлагаемая концепция призвана стать основой для раннего выявления гепатитов B и C, стандартизации мониторинга и, в перспективе, для разработки валидированных систем поддержки принятия клинических решений на основе машинного обучения.
Keywords: hepatitis B, hepatitis C, clinical-laboratory indicators, database, screening, early detection, Uzbekistan, digital medicine.
Ключевые слова: гепатит B, гепатит C, клинико-лабораторные показатели, база данных, скрининг, раннее выявление, Узбекистан, цифровая медицина.
Introduction
Hepatitis B and C viruses continue to be among the world's most urgent health issues. These diseases are particularly dangerous because they frequently have a latent or asymptomatic course. Because of this, the patient is typically diagnosed at a far more advanced stage of the illness. Since early infection can go undiagnosed for years and eventually show up as chronic liver damage, this problem is particularly significant in children and adolescents. Thus, in pediatric hepatology, early detection, routine monitoring, and risk group differentiation are critical tasks. This subject is particularly pertinent to Uzbekistan.A large-scale national screening program for hepatitis B and C has been launched in the country in 2022–2025, and more than 1 million people have been tested in the initial stage. According to the study, the overall prevalence of HBsAg was 2.89%, and the prevalence of anti-HCV was 3.52%. Significant differences were found across regions: the highest HBV rate was in Bukhara region at 3.96%, the lowest in Tashkent city at 2.14%; the highest HCV rate was in Tashkent city at 6.07%, the lowest in Jizzakh region at 1.29%[1]. This epidemiological picture shows the need for a structured, multi-parameter analytical approach that combines several parameters and not just a single laboratory result in the detection and monitoring of hepatitis.
The purpose of this article is to systematize the necessary clinical and laboratory indicators for the detection of hepatitis B and C, propose a database structure based on them, and develop a concept for a digital software module that can be used in the conditions of Uzbekistan.
Materials and methods
The article's theoretical and practical foundation came from Uzbekistan's national screening program, which ran from July 2022 to June 2024. 1,048,575 individuals between the ages of 1 and 95 took part in the screening, which was carried out in 14 regions. The first screening step involved the use of rapid HBsAg and anti-HCV tests. Additional confirmatory tests were carried out in cases of positive results: HCV RNA PCR tests were carried out in anti-HCV positive cases, and HBsAg tests were carried out in HBsAg positive cases [1]. We are able to organize the data from the standpoint of clinical and epidemiological management thanks to this study. For every patient, a single electronic card is used to create the database. It is filled with the block parameters listed below.
Table 1. Parameters of the electronic cards
|
№ |
Parameters |
|
1. |
Demographic information |
|
2. |
Clinical signs |
|
3. |
Laboratory indicators |
|
4. |
Virological data |
|
5. |
Instrumental examination results |
|
6. |
Final risk category |
In addition, the system automatically calculates derived indices based on clinical laboratory data:
- AST/ALT ratio;
- APRI index;
- FIB-4 index;
- log-transformation of viral load.
These indices are widely used, guideline-endorsed tools for the non-invasive assessment of the severity of hepatitis and liver fibrosis: the APRI index was introduced by Wai et al. [5] and the FIB-4 index by Sterling et al. [6]. In the present work they are not inputs to a trained statistical model; instead, their published cut-off values are applied directly as transparent, threshold-based (rule-based) criteria for assigning patients to risk groups.
Conceptual database design
It should be emphasized that the structure presented here is a conceptual (logical) relational design rather than a deployed and populated production system. The model is normalized into linked tables, each connected to a central patient table through a unique patient identifier (patient_id) used as the primary key and propagated as a foreign key to the block tables; the one-to-many relationship between the patient table and the measurement tables allows repeated (longitudinal) records to be stored for the same individual. A relational database management system such as PostgreSQL is proposed as the target implementation environment, with derived indices computed at the application layer. Table 2 Summarizes the conceptual schema
Table 2. Conceptual relational schema of the proposed database
|
Table (entity) |
Key fields and relationships |
|
patient |
patient_id (PK), registration_date, region; central table referenced by all others |
|
demographics |
patient_id (FK), age, sex, district |
|
clinical_signs |
patient_id (FK), symptom indicators, examination notes; one-to-many (longitudinal) |
|
laboratory_results |
patient_id (FK), ALT, AST, bilirubin, albumin, platelet count, INR, test_date |
|
virological_data |
patient_id (FK), HBsAg, anti-HCV, HBV DNA, HCV RNA (PCR) |
|
instrumental_findings |
patient_id (FK), ultrasound result, elastography (where available) |
|
derived_indices |
patient_id (FK), AST/ALT ratio, APRI, FIB-4, log(viral load); computed values |
|
risk_classification |
patient_id (FK), risk_category, applied_rule_version, timestamp |
Rule-based risk-stratification criteria
Risk groups are assigned by a deterministic, threshold-based rule set rather than by a trained classifier. The published cut-off values of the two indices are used directly: a FIB-4 value below 1.45 and an APRI value below 0.5 indicate a low likelihood of advanced fibrosis, whereas FIB-4 above 3.25 or APRI above 1.5 indicate a high likelihood; intermediate values define the moderate-risk group [5, 6]. An AST/ALT ratio greater than 1, together with a low platelet count and reduced albumin, is treated as a supporting indicator of more advanced disease. A patient is placed in the highest category triggered by any applicable rule. Because the criteria are explicit and reproducible, every classification can be traced back to the underlying values, which is the defining property of a rule-based clinical decision-support system as opposed to a machine-learning model.
Proposed software module and principle of operation
The proposed digital system works as a clinical-laboratory platform that automatically processes data entered by a doctor or laboratory worker. It works in the following steps:
- receiving patient information;
- data verification and cleaning;
- calculation of derived indices;
- issuing a risk assessment based on clinical-laboratory criteria;
- save the result on an electronic card and show it to the doctor.
This approach is particularly useful in less symptomatic forms of hepatitis, allowing the physician to quickly assess the overall risk profile without manual comparisons.
As a future direction, the database has been designed so that the same input fields (age, sex, ALT, AST, bilirubin, albumin, platelet count, INR, viral load, and ultrasound findings) could later feed a supervised machine-learning classifier such as XGBoost [7]. It must be emphasized, however, that no such model has been trained or validated in the present study; the machine-learning component is described only as a planned extension. Realizing it would require a labelled clinical dataset, a defined training and validation split, hyperparameter tuning, and the reporting of performance metrics such as AUC, sensitivity, and specificity. Until that work is completed, the operational risk assessment relies solely on the transparent rule-based criteria described above.
Results and discussion
The outcomes of the extensive screening carried out in Uzbekistan offer a crucial epidemiological foundation for the creation of a database. 30,309 and 36,896 of the 1,048,575 people who were screened tested positive for HBsAg and anti-HCV, respectively. Both markers were found at the same time in 3,684 cases. These results highlight the critical need for an extensive database of indicators for managing and detecting diseases.
Regional analysis showed significant differences in the prevalence of hepatitis B and C. Bukhara region was identified as the area with the highest HBV prevalence, while Tashkent city demonstrated the highest HCV rate. Therefore, it is advisable to include the patient’s place of residence as a separate field in the database and to take it into account in epidemiological analysis.
Age and sex factors are also of major importance. The fact that HBsAg prevalence was 0.64% and anti-HCV prevalence was 1.66% in the 1–20 age group indicates a relatively lower spread of the disease among children and adolescents, but this does not mean that monitoring is unnecessary. On the contrary, early detection in this group can help prevent long-term complications. The prevalence of HBsAg was 3.70% in males and 2.58% in females, while anti-HCV was 4.18% in males and 3.17% in females. These findings support considering sex as a prognostic indicator in the database.
The confirmatory stage after screening is also important for the database. Among 32,132 individuals who tested positive for anti-HCV, active infection was confirmed by HCV RNA PCR in 20,039 cases, or 62.4%. Of these, 19,286 attended hepatology consultations, and 18,327 were found eligible for DAA treatment. These data suggest that the software system should manage not only screening results, but also confirmatory testing and subsequent referral stages.
To illustrate the intended workflow of the proposed system, an illustrative example based on 120 patient records was prepared: 46 HBV cases, 28 HCV cases, and 46 cases under differential follow-up. When the rule-based criteria were applied, 39 records fell into the low-risk group, 51 into the moderate-risk group, and 30 into the high-risk group. In the high-risk group, ALT, AST, bilirubin, and INR tended to be higher, whereas albumin and platelet counts tended to be lower, which is consistent with the expected direction of these markers in more advanced liver disease. This example is descriptive and illustrative only: it was not designed as a validation study, no comparison against a clinical reference standard (ground truth) was performed, and no diagnostic-accuracy metrics such as sensitivity, specificity, or AUC are reported. It therefore demonstrates how the rule set assigns categories, but it does not establish the diagnostic performance of the system.
Beyond its immediate use, the standardized database also provides a foundation for future, machine-learning-based clinical decision-support systems. Risk stratification could in principle be learned by models such as XGBoost [7], building on prior machine-learning approaches to hepatitis detection and paediatric liver-fibrosis prediction [8–10], once the database is enhanced with age-specific laboratory reference values, standardized data-entry formats, and a sufficient volume of labelled clinical data. Such a model would, however, need to be properly trained and validated — with a defined training and test split, cross-validation, and reported metrics (AUC, accuracy, sensitivity, and specificity) — before any machine-learning claim could be made.
This study has several limitations that should be stated explicitly. The database is presented as a conceptual (logical) design and has not yet been implemented as a deployed system or populated for production use. The risk-stratification component is rule-based and uses published index cut-offs; it has not been calibrated to age-specific paediatric reference ranges. No machine-learning model has been trained or validated, and the 120-record example is illustrative rather than a diagnostic-accuracy study. Prospective evaluation against clinical reference standards, multi-centre testing, and local validation are required before the framework can be recommended for routine clinical use.
Conclusion
The formation of clinical-laboratory indicators and a database for the detection of hepatitis B and C is a relevant and practically significant direction in the context of Uzbekistan. The nation has an adequate epidemiological foundation for this work, as evidenced by the national screening program carried out in 2022–2024: wide population coverage, regional variations, age and sex factors, and notable variations at the confirmatory stage all support the necessity of gathering data in a standardized electronic format. The proposed conceptual database combines the demographic, clinical, laboratory, virological, and instrumental indicators into a single electronic record, and the accompanying rule-based criteria use established non-invasive indices to assign risk groups and direct patients for additional testing. While the system has not yet been implemented or validated, this methodological foundation can support the standardization of monitoring, the early detection of hepatitis, and, as a future direction, the development of validated machine-learning-based clinical decision-support systems.
References:
- Ismoilov UY, Musabaev EI, Khikmatullaeva AS, et al. Hepatitis B and C landscape in Uzbekistan: epidemiological patterns revealed in a study of 1 040 000 people. European Journal of Public Health. 2025;35(5):1014–1019. doi:10.1093/eurpub/ckaf136.
- World Health Organization. Global hepatitis report 2024: action for access in low- and middle-income countries. Geneva: World Health Organization; 2024.
- World Health Organization. Guidelines for the prevention, diagnosis, care and treatment for people with chronic hepatitis B infection. Geneva: World Health Organization; 2024.
- Wong GL, Lemoine M. The 2024 updated WHO guidelines for the prevention and management of chronic hepatitis B: main changes and potential implications for the next major liver society clinical practice guidelines. Journal of Hepatology. 2025;82(5):918–925. doi:10.1016/j.jhep.2024.12.004.
- Wai CT, Greenson JK, Fontana RJ, Kalbfleisch JD, Marrero JA, Conjeevaram HS, Lok AS. A simple noninvasive index can predict both significant fibrosis and cirrhosis in patients with chronic hepatitis C. Hepatology. 2003;38(2):518–526. doi:10.1053/jhep.2003.50346.
- Sterling RK, Lissen E, Clumeck N, Sola R, Correa MC, Montaner J, et al. Development of a simple noninvasive index to predict significant fibrosis in patients with HIV/HCV coinfection. Hepatology. 2006;43(6):1317–1325.
- Chen T, Guestrin C. XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. New York: ACM; 2016. p. 785–794. doi:10.1145/2939672.2939785.
- Li S, Li S, et al. Developing and validating a prediction tool for identifying significant liver fibrosis in children with chronic hepatitis B: an interpretable machine learning model. Computers in Biology and Medicine. 2025;199:111293. doi:10.1016/j.compbiomed.2025.111293.
- Obaido G, Ogbuokiri B, Swart TG, et al. An interpretable machine learning approach for hepatitis B diagnosis. Applied Sciences. 2022;12(21):11127.
- Ahosan AB, Islam F, Mohi Uddin KM, Hasan N, Uddin MA. Improving hepatitis B outcome prediction with ensemble machine learning: a study on predictive models and interpretability. Digital Health. 2025. doi:10.1177/20552076251350755.