The project
Spanish doesn’t only live where it has always been spoken. It lives on the signage of a Berlin restaurant, in the classroom where someone learns German from Spanish, in the expressions two languages borrow from each other without asking.
Hispanoverso/DemIA gathers the research documenting that Spanish in motion and turns it into something you can actually consult: an open catalogue and a search that understands questions in plain language.
Where it comes from
The Hispanoverso research group at the University of Salamanca works on digital communication and the global projection of Hispanic culture. This project is its data infrastructure: it puts into circulation what the group has published over a decade.
It also sits within the International Chair in Trustworthy Artificial Intelligence and the Demographic Challenge (DemIA), which addresses depopulation, ageing and territorial inequality through explainable AI, from an ethical, people-centred perspective.
The connection between the two isn’t incidental. Migration is the other face of the demographic challenge: emptying territories and Spanish-speaking communities forming abroad are the same phenomenon seen from two shores. And language is where that movement leaves its first trace.
What it holds
144 publications spanning 2017 to 2026, in Spanish, German and English.
The corpus covers language contact and language teaching, the linguistic landscape of European cities, contrastive phraseology, shared cultural heritage and — increasingly — the responsible use of artificial intelligence in training language teachers.
Chapters in edited volumes predominate, followed by journal articles. This is the characteristic profile of philological research, and it has a practical consequence: most of these publications don’t appear in the databases that assign subject descriptors, and fewer than a third have a DOI.
Put another way: it is valuable research that is hard to find. That gap is precisely what this project sets out to close.
How it works
Each night the system queries the University of Salamanca research portal and collects each publication’s metadata in BibTeX, a bibliographic interchange standard. Compared with scraping web pages, this yields already-structured data and survives redesigns of the portal.
Where a publication has a DOI, it is completed from OpenAlex: abstract, open access availability, a link to the document and a citation count.
Each publication is then turned into a numerical representation of its content. That is what allows searching by meaning: someone can type “how is technical vocabulary taught in language classes?” and find relevant work even when it shares no words with the question.
The catalogue is published as static pages. With no database queried on each visit, it loads in milliseconds and search engines index it without friction.
How it is organised
Publications are distributed across six thematic axes: languages in contact, linguistic landscape, language teaching, migration and diaspora, cultural heritage, and AI and language.
Assignment uses explicit rules over the title, the containing publication and keywords where present. The criterion is deliberately auditable: anyone can examine the rules and understand why a publication appears under a given axis.
Part of the corpus remains unclassified. This follows directly from the scarcity of descriptive metadata at source, and is documented openly rather than forced into a category it doesn’t belong to.
Who it is for
For researchers who need to know what has already been published on a specific question.
For teachers looking for materials and references on language pedagogy.
For evaluators who want to see, at a glance, what a research group produces and how it evolves.
And for anyone curious about how Spanish travels and what happens to it when it meets other languages.
Open data
Bibliographic metadata is published under CC0 and accessible through a documented public interface. Original documents retain their publishers’ licences.
The approach matches that of the DemIA Living Lab: a space where researchers, institutions and the public can use, experiment with and download digital resources openly and free of charge.
The system’s source code is publicly available.