Multilingual corpora for the study of new concepts in the social sciences and humanities:

View PDF

Abstract:This article presents a hybrid methodology for building a multilingual corpus designed to support the study of emerging concepts in the humanities and social sciences (HSS), illustrated here through the case of ``non-technological innovation’’. The corpus relies on two complementary sources: (1) textual content automatically extracted from company websites, cleaned for French and English, and (2) annual reports collected and automatically filtered according to documentary criteria (year, format, duplication). The processing pipeline includes automatic language detection, filtering of non-relevant content, extraction of relevant segments, and enrichment with structural metadata. From this initial corpus, a derived dataset…

View PDF

Abstract:This article presents a hybrid methodology for building a multilingual corpus designed to support the study of emerging concepts in the humanities and social sciences (HSS), illustrated here through the case of ``non-technological innovation’’. The corpus relies on two complementary sources: (1) textual content automatically extracted from company websites, cleaned for French and English, and (2) annual reports collected and automatically filtered according to documentary criteria (year, format, duplication). The processing pipeline includes automatic language detection, filtering of non-relevant content, extraction of relevant segments, and enrichment with structural metadata. From this initial corpus, a derived dataset in English is created for machine learning purposes. For each occurrence of a term from the expert lexicon, a contextual block of five sentences is extracted (two preceding and two following the sentence containing the term). Each occurrence is annotated with the thematic category associated with the term, enabling the construction of data suitable for supervised classification tasks. This approach results in a reproducible and extensible resource, suitable both for analyzing lexical variability around emerging concepts and for generating datasets dedicated to natural language processing applications.


Comments:	in French language
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2512.07367 [cs.CL]
	(or arXiv:2512.07367v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2512.07367 arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Anna Pappa [view email] [via CCSD proxy] [v1] Mon, 8 Dec 2025 10:04:50 UTC (419 KB)

Submission history

Similar Posts