Abstract
A corpus of religious texts is a system of processed religious texts whose creation improves the understanding of textual content, automatic semantic analysis, and machine translation. Based on the publication of Islamic scholar Muhammad Sodiq Muhammad Yusuf's "Sahih al-Bukhari" vol. 1 (2019), the article presents the semantic field and semantic groups among 428 studied lexical units and analyzes the units "Book" and "Hadith" using corpus-linguistic methods of thematic grouping, semantic tagging, and keyword analysis, resulting in 17 thematic groups.