Abstract:
This article presents an additive Multidimensional (MD) analysis (Biber, 1998) of the LEX-BR-Ius (Ferrari; Marques, 2022), a corpus of Brazilian federal statutory laws. We first contextualize our research and, in sequence, introduce the MD approach. Subsequently, we present the corpus under study and the Brazilian Register Variation Corpus (Corpus Brasileiro de Variação de Registro - CBVR) (Berber Sardinha; Kauffmann; Acunzo, 2014), to which our corpus has been added, along with a step-by-step description of the methodology adopted. Lastly, we present and discuss the results obtained, the issues that emerged during the research, and their potential explanations.
Keywords:
Multidimensional analysis; LEX-BR-Ius; corpus; methodology; register; legal language
Resumo:
Este artigo apresenta a Análise Multidimensional Aditiva (AMD) (Biber, 1988) do LEX-BR-Ius (Ferrari; Marques, 2022), um corpus de leis federais brasileiras. Primeiramente, contextualizamos nossa pesquisa e, na sequência, introduzimos a abordagem da AMD. Em seguida, descrevemos o corpus em estudo e o Corpus Brasileiro de Variação de Registro (Berber Sardinha; Kauffmann; Acunzo, 2014), ao qual nosso corpus foi adicionado, além de um passo a passo da metodologia adotada. Por fim, apresentamos e discutimos os resultados obtidos, os problemas que surgiram durante a pesquisa e suas possíveis explicações.
Palavras-chave:
Análise Multidimensional; LEX-BR-Ius; corpus; metodologia; registro; linguagem jurídica
1 Introduction
The Multidimensional analysis (MD) approach (Biber, 1988) has been considered for years as a powerful tool to describe the linguistic variation among registers (Biber; Conrad, 2009). Since introduced by Biber (1988), numerous studies have applied this method to examine register variation and identify co-occurrent linguistic features associated with different registers and languages (Berber Sardinha, 2010; Delfino, 2021; Berber Sardinha; Kauffmann; Acunzo, 2014). In the MD analysis, a register is “[…] a variety associated with a particular situation of use (including particular communicative purposes)” (Biber; Conrad, 2009, p. 6). In other words, a register is a variety of language used in specific communicative contexts. It is composed of linguistic (lexico-grammatical), situational (context of production, participants, communicative purpose, communication channel, subject, etc.), and functional (describing the relationship between the first two) features allowing the description of a given register in its entirety (Biber; Conrad, 2009).
Within this framework, our study aims to identify the linguistic features associated with a distinct register, legislation, in Brazilian Portuguese (BP). This register belongs to a major domain called legal language, a specialized language used in a particular but also broad context that includes written documents such as legislation, court documents, contracts and administrative regulations, and wills as well as spoken interactions like courtroom proceedings, depositions, police interactions, and arbitrations (Tiersma, 1999; Goźdź-Roszkowski, 2012; Carapinha, 2018).
Usually deemed precise, highly formal, and grammatically complex, this language is characterized by technical lexicon, often with the use of archaisms and Latin vocabulary, nominalizations, passive voice, modal verbs that indicate mandatory actions or permissions, and phrases that impose obligations, rights, or prohibitions (Tiersma, 1999; Dahlman, 2006; Rossini Favretti; Tamburini; Martelli, 2001; Goźdź-Roszkowski, 2012; Chovanec, 2013; Goźdź-Roszkowski, Pontrandolfo, 2014; Svobodová, 2017; Richard, 2018; Lorz, 2019). Regardless, to the best of our knowledge, there are no previous studies that linguistically outline the peculiarities of Brazilian laws, and the MD approach is a promising method to trace this register’s complete and detailed profile.
This method was brought to Brazil by Berber Sardinha in the 1990s and, since then, he and his team at Linguística Aplicada e Estudos da Linguagem (LAEL) (PUC-SP) have applied this approach in a series of studies in different languages such as English, Portuguese, and Spanish (e.g. Silveira, 1997; Conde, 2002; Bértoli-Dutra, 2010; Berber Sardinha, 2010, 2022, 2024; Zuppardi, 2021; Delfino, 2021; Veiga, 2021; Escarabelin, 2022; Delfino; Berber Sardinha; Collentine, 2023; Braz, 2023; Berber Sardinha et al., 2022).
The researcher also created and improved the post-processor PT Tag Count (Berber Sardinha, 2013a), which tags and counts BP corpora in the 196 linguistic features considered relevant for the study of linguistic variation in this language (Berber Sardinha; Kauffmann; Acunzo, 2014). By performing an MD analysis, linguists can achieve a broad and detailed description of the lexical and grammatical features of the texts, and by functionally interpreting the results, can obtain a faithful portrait of the registers under analysis.
Encouraged by previous MD research on Brazilian Portuguese (BP) (Oliveira, 1997; Santos, 2003; Berber Sardinha, 2003; Berber Sardinha; Kauffmann; Acunzo, 2014; Berber Sardinha et al., 2019; Kauffmann, 2020), our objective is to describe thoroughly the language employed in Brazilian federal statutory laws. Additionally, we aim to verify the MD consistency and its possible issues. In this paper, we will present the additive MD analysis we performed on the LEX-BR-Ius corpus (Ferrari; Marques, 2022), discuss its results, highlight the issues faced when performing the analysis, as well as comment on some intriguing findings. The latter were re-investigated by scrutinizing the tags applied by the parser and the examples of each feature generated by the post-processor. We believe that this work can contribute to the description of a very specific register, i.e. the language used in federal statutory laws in Brazil, as well as present issues within the MD performed, pointing to some problems whose solution would improve this language portrayal.
2 MD Analysis: An Overview
Considered “one of the most productive analytical approaches used to describe the overall patterns of register variation in a discourse domain” (Biber; Conrad, 2009, p. 246), the Multidimensional (MD) analysis (Biber, 1988) is a corpus-based methodological approach to study linguistic variation. It is applicable to any language - such as English (Biber, 1988; Cao; Xiao, 2013), Korean (Kim; Biber, 1994), Gaelic (Lamb, 2002, 2008), Somali (Biber; Hared, 1994a, 1994b), Spanish (Parodi, 2007), and Portuguese (Berber Sardinha; Kauffmann; Acunzo, 2014) -, and to both synchronic (e.g. Biber, 1988, 1995) and diachronic studies (e.g. Biber; Finegan, 1989, 1992, 1997; Geisler, 2000, 2001a, 2001b; Berber Sardinha, 2010; Kauffmann, 2020). Idealized by Douglas Biber in the 1980s, this approach revolutionized the form language was analyzed using large corpora, inaugurating the American approach to Corpus Linguistics (Delfino, 2021).
The MD analysis comprehensively describes linguistic variation within and across multiple registers and sub-registers in dimensions of variation (Biber, 1988), achieved through the statistical analysis of several linguistic features, from which co-occurrence patterns - factors - are thus identified. These patterns, functionally interpreted, reveal underlying dimensions of variation (Berber Sardinha, 2010; Biber; Conrad, 2009). By comparing registers and sub-registers in these dimensions it is possible to pinpoint their linguistic features, similarities, and differences regarding their textual constitution and internal organization (Delfino, 2021).
As stated by Biber (1988):
This approach is based on the assumption that strong co-occurrence patterns of linguistic features mark underlying functional dimensions. Features do not randomly co-occur in texts. If certain features consistently co-occur, then it is reasonable to look for an underlying functional influence that encourages their use. In this way, the functions are not posited on an a priori basis; rather they are required to account for the observed co-occurrence patterns among linguistic features (Biber, 1988, p. 13).
In other words, a set of linguistic feature patterns co-occurs in a given register to fit the register’s communicative purposes and situational context and, thus, it is functionally motivated. The functional interpretation of these patterns originates the dimensions of variation. Biber (1988) defends that a language cannot be adequately described in a single dimension: to describe a language in its entirety it is necessary to analyze it under multiple dimensions to encompass its many aspects.
Each dimension consists of a distinct set of co-occurring linguistic features, which are quantitatively identified through factor analysis. These features establish a continuum of variation, where registers are distributed according to their dimension scores in scales. Each dimension typically consists of a scale with two opposing poles, negative and positive, in complementary distribution. The registers are distributed according to their dimension scores in these scales. The closer the register is to one of the poles, the more prototypical it is of the pole. By examining the registers in all dimensions, we can describe them and understand the overall patterns of register variation in each dimension (Biber, 1988; Berber Sardinha, 2010; Delfino, 2021).
To perform the MD analysis, first and foremost, we need a corpus of the registers we want to analyze. Making an accurate description of a given register and establishing generalizations regarding it requires a balanced sample of complete authentic texts representative of the register compiled in electronic format (Biber, 1988; Biber; Conrad, 2009; Brezina, 2018). Ideally, each register sampled in the corpus must have a similar number of texts, words, and words per text (Brezina, 2018). Additionally, the corpus must be annotated according to the linguistic features that are going to be studied, established based on previous studies about the language variety under investigation (Biber, 1988; Berber Sardinha, 2010; Brezina, 2018; Delfino, 2021).
The next step is to count the linguistic features in the corpus and normalize their frequency to establish a common ground in size. The computed tables of frequency are then subjected to factor analysis, a statistical procedure that identifies groups of co-occurrent variables (the linguistic features) in the corpus. Each group represents a factor, and its stronger forming features are accounted for every text in the corpus based on their standardized counts, resulting in individual factor scores. The registers’ dimension scores are then calculated from mean dimension scores derived from register sub-corpora. Lastly, the results are functionally and discursively interpreted, establishing the dimensions of variation in which the registers are arranged on a scale according to their factor/dimensions’ scores and compared to each other (Biber, 1988; Berber Sardinha, 2010; Brezina, 2018; Delfino, 2021).
When performing the MD analysis, we can choose between the full MD analysis and the additive MD analysis. The first follows all the steps proposed by Biber (1988) and described above, identifying the dimensions of variation through factor analysis and providing a detailed picture of the registers analyzed in these dimensions. In the second modality, instead of identifying the dimensions of variation, the dimensions identified in an existing full MD analysis, referred to here as the base study, are used, making it unnecessary to carry out factor analysis (Berber Sardinha et al., 2019). Additionally, as it is with the full MD analysis, the additive MD analysis also allows the realization of both synchronic (e.g. Biber, 1987; Conrad, 1996, 2001, 2014; Biber et al., 2002; Helt, 2001; Rey, 2001; Connor; Upton, 2003; Forchini, 2012; Jonsson, 2016) and diachronic studies (e.g. Biber; Finegan, 1988, 2001; Atkinson, 1992, 1996, 2001; Biber; Hared, 1992a, 1992b; Souza, 2014; Veirano Pinto, 2014). Due to that, as pointed out by Berber Sardinha et al. (2019), the additive MD analysis is known for its flexibility, technical simplicity, and shorter time required to perform it when compared to the full MD analysis.
As highlighted by Delfino (2021), to perform the additive MD analysis, the corpus under study must be collected, tagged, and analyzed in the same way as the corpus of the base study. Furthermore, the language of the register(s) under analysis must be the same as that of the base study. Finally, the frequencies of the linguistic variables of the corpus under study must be standardized according to the mean and standard deviation of the base study. This step is done by adding and comparing the registers under study to the base study dimensions (Berber Sardinha et al., 2019).
As Berber Sardinha et al. (2019) emphasize, full and additive MD analyses are complementary approaches to studying register linguistic variation. While the former establishes the dimensions of variation of a given language or variety, the latter adds new registers to a previous full MD study, increasing its scope and enriching the variation description.
Regarding previous MD research on our target language, Brazilian Portuguese, to the best of our knowledge, the only studies are: Oliveira (1997), Santos (2003), Berber Sardinha (2003), Berber Sardinha, Kauffmann, and Acunzo (2014), Kauffmann (2020), and Kauffmann (2005). Oliveira (1997) described written production from students in BP and English. On the other hand, Santos (2003) focused her research on the register business manual, mapping its keywords onto Biber’s (1988) dimensions. With a similar approach, Berber Sardinha (2003) researched linguistic variation in a business meeting corpus, also extracting its keywords and adding them to Biber’s (1988) dimensions. Conversely, Kauffmann (2005) studied linguistic variation in news registers in a famous Brazilian newspaper.
Kauffmann (2020), in turn, investigated the linguistic variation in Machado de Assis’s prose from the style perspective proposed by Biber and Conrad (2009). Interestingly, Kauffmann (2020) performed multiple MD analyses in his study. He performed two additive ones, the first comparing Machado de Assis work with the work of other contemporary writers of his, and the second mapping his Machado de Assis corpus in the Brazilian Portuguese dimensions of variation. In the sequence, three full MD analyses were also performed, two using functional tags and one from a lexical perspective, using lemmas as tags. Lastly, and differently from all the other BP MD studies, Berber Sardinha, Kauffmann, and Acunzo (2014) aimed to describe Brazilian Portuguese variation in its entirety, performing the most comprehensive full Multidimensional analysis ever carried out in this language. This analysis comprised 48 different registers and sought to draw the most detailed portrait of BP possible. As a result, they identified 6 BP dimensions of variation that we will further describe in Section 3.2.
In summary, all the studies regarding BP adopted the synchronic perspective and, while Oliveira (1997), Berber Sardinha, Kauffmann, and Acunzo (2014), and Kauffmann (2005) are full MD analyses, Santos (2003) and Berber Sardinha (2003) are additive ones. Kauffmann (2020), by its turn, performed both full and additive MDs in his study. Except for Berber Sardinha, Kauffmann, and Acunzo (2014), the only one that aimed to describe the overall variation of BP, the others were focused on specific registers, namely student writings, newspapers, manuals, and business meetings.
It is important to notice that using an MD analysis of a language different from the one being researched as the base study for an additive MD, as in Santos (2003) and Berber Sardinha (2003), is not recommended. This approach may lead to methodological issues because different languages often have different linguistic features and underlying dimensions, which may compromise the comparability of the data. Applying linguistic patterns identified in one language to another can lead to skewed results due to the structural and functional differences between the languages involved in the analysis (Berber Sardinha, 2010; Delfino, 2021).
3 Study Design and Data Analysis
In this section, we will address the corpora used in this study, namely the LEX-BR-Ius (Ferrari; Marques, 2022) and the CBVR (Berber Sardinha; Kauffmann; Acunzo, 2014). While the first is described in section 3.1, the latter is presented in section 3.2. Section 3.3 exhibits a detailed description of the additive MD analysis performed in the study.
3.1 The LEX-BR-Ius Corpus
The LEX-BR-Ius (Ferrari; Marques, 2022) is a corpus of Brazilian federal statutory laws in force on the compilation date. The corpus includes 755 different laws, collected in their entirety, totalizing 3,300,289 tokens.
To the best of our knowledge, this is the first representative corpus of Brazilian laws. The closest work we found is a dataset by Martim, Lima, and Araújo (2018), however, its architecture and purposes diverge from ours. Following the pillars of corpus linguistics, we propose that legislative data should be sampled in a representative way, and special attention is given to the texts cleansing and annotation. In Brazil, laws are created and modified daily, in all three spheres of the government: municipal, state, and federal.
For this corpus, federal statutory laws were chosen for their national scope. Before starting the data collection, all federal statutory laws available at Portal da Legislação (http://www4.planalto.gov.br/legislacao/) - a governmental website that stores Brazilian law texts, which are constantly updated - were verified and listed. We aimed to collect only laws in force during the compilation period (January 2022 to January 2024). To do so, our team checked this information, discarding the laws that were no longer in effect. It is important to elucidate that some laws in force today date from as far as 1850 (e.g. Código commercial - Commercial Code), long before our actual Constitution, of 1988, came into force, and so, some of these texts were massively modified throughout the years. Meanwhile, other legal texts, although still in force, present little modifications.
To build our corpus, we first downloaded all the laws previously selected from the Portal da Legislação in a .docx format to preserve the different textual information; for example, the previous versions of parts of the law that are strikethrough in the text to indicate they were no longer in effect. For each law, the .docx document was kept in the folder as the original, and a copy was converted into .txt format in the Notepad ++ editor (Ho, 2020). The texts were manually cleaned, excluding metadata, hyperlinks, figures, page numbers, and tables. This raw text was then annotated using the PALAVRAS tagger (Bick, 2000, 2014), considered the state-of-the-art tagger for Brazilian Portuguese, to allow for deep linguistic research as well as the realization of the MD analysis. A metadata header was also compiled with information such as dates of creation and promulgation of the law, the president in charge at the time, its modifications, number of words etc. A third version of each text was made with manual textual markup following Hardie’s (2014) Modest XML proposal, with tags specifically created for this corpus.
To select the laws that would be included in our corpus, we adopted specific criteria of representativeness: the usage of each law in the legal and judicial world (Barbera; Onesti, 2009; Onesti, 2011). Additionally, we searched to include laws from all types in effect available in the Portal da Legislação, namely: Códigos (Codes), Constituição (Constitution), Emendas à Constituição (Amendments to the Constitution), Estatutos (Statutes), Leis Complementares (Complementary laws), and Leis Ordinárias (Ordinary laws). Their use rate was verified in the Jusbrasil portal (https://www.jusbrasil.com.br/), the biggest aggregator website of Brazilian law-related content, selecting the ones with over 10,000 citations. Table 1 shows this sampling in detail.
As shown in Table 1, our sample was divided into six sections, following the Portal da Legislação classification. We selected 13 out of 17 Códigos and 11 out of 18 Estatutos. They were selected for their importance in the legal world and we excluded those aimed at very restricted groups, such as the Código Militar (Military Code) and the Estatuto dos Museus (Museum Statute). The major issue was sampling Emendas à Constituição, Leis Complementares, and Leis Ordinárias. Due to their often limited and over-specific scope, only a small subset of these laws met our citation criteria, respectively 48 out of 114, 43 out of 192, and 639 out of 13,491 laws.
Besides the representativeness of our sample, another methodological issue refers to the corpus balancing. Different theories and methods suggest what an ideal balancing should be, going from a similar number of words to excerpts of 600 to 1,000 or 2,000 words per text (Biber, 1993; Atkins; Clear, 1992; Sinclair, 2004; McEnery; Xiao; Tono, 2006; Leech, 2007; McEnery; Hardie, 2012). Laws vary considerably in length, from a few tens (e.g. EC 90_15.09.2015, with 56 words) to hundreds of thousands of words (e.g. LO10.406_10.01.2002, with 101,629 words). Grouping these texts based on their number of words would be a possibility, but it would skew the data, so, following Sinclair (2004) and Biber (1993), in order to preserve the laws’ internal organization, textual variation, representativeness, and structure, we decided to maintain their integrity, using only complete texts (Ferrari; Marques, 2022).
3.2 The CBVR
The reference corpus used in this study, Brazilian Register Variation Corpus (Corpus Brasileiro de Variação de Registro - CBVR) (Berber Sardinha; Kauffmann; Acunzo, 2014), was designed to represent contemporary Brazilian Portuguese language in use. It has a large coverage, from traditional, literary varieties, such as academic articles, general fiction, and TV news, to more prosaic texts, e.g. jokes, user manuals, and posts from social media, comprising 12 oral and 36 written registers. The CBVR has 960 texts - 20 per register - reaching 5.64 million words in total.
Seven groups of registers could be devised (Berber Sardinha; Kauffmann; Acunzo, 2014): media, academic-pedagogic, literary, public-administrative, online, oral, and other registers. The public-administrative sphere of CBVR’s (Berber Sardinha; Kauffmann; Acunzo, 2014) registers includes Minutes, Business letters, Agreements, Legislation, Government bids, and Campaign plans (written), as well as Political speeches, Congressional debates, and Business conference calls (oral). Public-administrative registers constitute a literate group, marked also with non-argumentative and non-directive discourses (Berber Sardinha; Kauffmann; Acunzo, 2014). The sample of 20 texts gathered to represent the legislation register in the corpus include Federal Laws (12), Codes (3), Interim orders (2), Additional acts (2), and part of the Constitution of 1988 itself, in one file with its first 20,000 words.
The six dimensions of variation identified by Berber Sardinha, Kauffmann, and Acunzo (2014) established important parameters for mapping previously included and new Brazilian Portuguese registers along the MD framework. Register mean scores computed on each dimension could be plotted along the dimension axis, providing ways for a comparable and a more detailed register description, as they show unique profiles across dimensions. The dimensions were labeled as:
-
(1) Oral versus Literate Discourse
-
(2) Argumentation
-
(3) Involved versus Informational Production
-
(4) Directive discourse
-
(5) Future versus Past Orientation
-
(6) Reported Discourse
Dimensions 1 and 3 have some intercorrelation and registers are more or less distributed in the same pattern, grouping oral and involved registers in one side and literate and informative registers in the other. Argumentation is expressed on Dimension 2 by opinion-driven registers as Political speeches and Interviews, but also Horoscope. Dimension 4 has a single pole linked to instructional registers, like user Manuals and Recipes, whereas Dimension 5 is occupied in separate registers rooted in past tenses from registers in which future tenses are relatively predominant. Finally, in Dimension 6 there is a concentration of literary, news and formal registers in its higher position, indicating frequent use of reported speech resources.
3.3 The Additive MD Analysis
As shown in section 2, Multidimensional (MD) analysis has been widely acknowledged as a highly effective approach for describing register linguistic variation (Biber; Conrad, 2009). It allows for the identification of registers’ lexico-grammatical patterns and their respective underlying functions in discourse through their placement and comparison in the established dimensions of variation (Biber, 1988). The dimensions are identified independently of register categories, thus enabling to study in detail the variation among any register or sub-register in the target corpus (Biber; Conrad, 2009). Due to that, we decided to apply this methodology to our corpus. As a full MD analysis of Brazilian Portuguese already exists (section 3.2), and given its advantages (i.e. simplicity and flexibility), performing an additive MD analysis has proven to be the ideal option to enable our research.
To perform it, we applied this methodology following Biber (1988), Berber Sardinha (2013b), and Berber Sardinha, Kauffmann, and Acunzo (2014), and using the software: SAS (SAS Institute, 2022), Microsoft Excel 2016, Notepad++ 7.8.9 (Ho, 2020), PALAVRAS (Bick, 2000, 2014), and PALAVRAS Tag Count 4 (Berber Sardinha, 2013a).
In this section, we will present the steps taken, summarized in Figure 1.
First, we mapped MD studies in BP and examined studies on legal language and the language used in laws, especially those that described its linguistic and situational features.
We then chose the study by Berber Sardinha, Kauffmann, and Acunzo (2014), here referred to as the base study, on the linguistic variation of Brazilian Portuguese to add the LEX-BR-Ius corpus (Ferrari; Marques, 2022). As shown in section 3.2, the researchers carried out a complete MD analysis of the Brazilian Register Variation Corpus (CBVR) (Berber Sardinha; Kauffmann; Acunzo, 2014), identifying the dimensions of variation in Brazilian Portuguese, which were applied to our corpus.
Having chosen our base study, the next step was to individualize the linguistic features of each of the identified BP dimensions. To make the additive MD analysis possible, we tagged our corpus in Part-of-Speech (PoS) and annotated it according to the linguistic features individualized in the previous step. To do this, we used the same software used in the base study: the PALAVRAS tagger (Bick, 2000, 2014) and post-processor PT Tag Count (v. 4) (Berber Sardinha, 2013a). The latter is a script that adds specific linguistic tags proposed by Biber (1988), such as abstract nouns and adverbs hedge, and automatically counts the annotated linguistic features, resulting in a “.csv” spreadsheet with the raw frequency of each feature in each text in the corpus. It is important to clarify that for unknown reasons the PT Tag Count (v. 4) (Berber Sardinha, 2013a) did not process one of our texts (LO11.496), so we excluded it from the MD analysis, applying it to 754 texts instead of the 755 of our corpus.
Afterward, following Biber (1988) and Berber Sardinha, Kauffmann, and Acunzo (2014), we normalized the number of occurrences of each feature by 1,000 words so that texts with different sizes could be equally compared. Additionally, we calculated their Z-scores to balance the features’ normalized frequencies so that they all have the same weight, preventing their values from influencing the next steps. This calculation was done using the normalized frequency of the features calculated in the previous step and the average frequency and standard deviation of the features performed in the base study. To do this, we used the formula (1):
With the Z-scores, we calculated the dimension scores for each text in the corpus. This calculation is used to determine the score, also called factor loading, of each text in our corpus in each dimension analyzed in the base study. To do this, we used the formula (2):
For dimensions with only one pole, the Z-scores values were simply added up.
After calculating the dimension score per text, the next step was to compute the average dimension score of the corpus and sub-corpora to map them into each dimension’s scale. This step allowed us to compare our corpus with registers of the base study in terms of dimensions. This step was carried out with the formulas (3) and (4):
To allow the comparison between our corpus and the registers of the base study, we calculated the dimension averages of the LEX-BR-Ius corpus (Ferrari; Marques, 2022) in its entirety and by section and added them to a spreadsheet with the dimension averages of the base study.
Lastly, to measure the range of variation covered by our corpus and its sub-corpora in each BP dimension and to determine if there are significant differences between them and between the base study registers, we performed two ANOVA tests (F-test, p-value, and R²) using SAS (SAS Institute, 2022). While the first test was conducted solely with our sub-corpora, the second compared a sample of our corpus with the CBVR registers. The sampling is required to ensure the comparability between the corpora. It comprised 20 texts from all of our sub-corpora randomly selected.
These tests allow us to measure the likelihood of the variation of each register in each dimension, one complementing the other. While the F-test and p-value identify whether there are significant differences between the register in each dimension, the R² test measures in percentage how much of this variation can be explained by the texts in each register (Berber Sardinha, 2013b; Brezina, 2018).
4 Results
In this section, we will present the results of the additive Multidimensional analysis performed by adding the LEX-BR-Ius (Ferrari; Marques, 2022) and its sub-corpora to the Brazilian Portuguese dimensions of variation identified by Berber Sardinha, Kauffmann, and Acunzo (2014). For each of the six dimensions, we bring its corresponding scale of variation in graph format, its linguistic features, and some examples from our corpus to illustrate the use of the aforementioned features.
It is important to highlight that in some dimensions our corpus and sub-corpora positions in the dimension’s scale differed from what we expected. To understand these findings, we examined all the tagged files checking for possible errors in the PALAVRAS tagger (Bick, 2001, 2014) and in the PT Tag Count (Berber Sardinha, 2013a). Our discoveries will be pointed out in this section and further discussed in Section 5.
4.1 Dimension 1: Oral Versus Literate Discourse
This first dimension characterizes the registers in terms of oral and literate discourses. The CBVR (Berber Sardinha; Kauffmann; Acunzo, 2014) registers and the LEX-BR-Ius (Ferrari; Marques, 2022) corpus and sub-corpora (in red in the Figure) are displayed in Figure 2 according to their dimension scores. While the registers on the positive pole present a higher occurrence of oral discourse features, those on the negative pole have a predominance of literate discourse features (see Table 2 for the features roll). The closer a register is to a pole, the more frequently its features occur in its texts; the further away, the more they are absent.
As shown in Figure 2, the most prototypical register of the positive pole in this dimension is Songs, while the most prototypical register of the negative pole is our sub-corpus Emendas à Constituição. Additionally, our corpus as a whole, the Leis Ordinárias, and Emendas à Constituição sub-corpora have extremely high scores. Interestingly, despite its similarities with the LEX-BR-Ius (Ferrari; Marques, 2022), the register Legislation (in green) has a much lower score than our corpus and is closer to our Códigos sub-corpora.
Our texts contain an extensive number of reduced gerund clauses (e.g. i), agentless passives (e.g. ii, iii, iv and v), compound names (e.g. ii and v), subject position nominalizations (e.g. iii), “rare” relative pronouns, such as qual and cujo (e.g. iv), topical adjectives (e.g. vi), attributive adjectives (e.g. v), and relational adjectives (e.g. vii), in addition to abstract nouns (e.g. v), and past participles (e.g. vi) These features reflect a highly formal, literate, specialized, planned discourse with a high informational density. Our corpus also presented a higher average size of the words when compared to the other registers. Furthermore, as legal texts require an elevated degree of accuracy, the frequency of definite articles and prepositions is significantly higher.
-
(i)I - praticar ato visando (reduced gerund clause) fim proibido em lei ou regulamento ou diverso daquele previsto, na regra de competência; (LO8.429_02.06.1992)
-
(ii)Art. 41. É suspensa (agentless passive) a execução da pena de multa (compound noun), se sobrevém ao condenado doença mental (compound noun). (C2.848_07.12.1940)
-
(iii)Art. 15. A participação (nominalization in subject position) no CONARE será considerada (agentless passive) serviço relevante e não implicará remuneração de qualquer natureza ou espécie. (E9.474_22.12.1997)
-
(iv)Art. 102. Sobrevindo o trânsito em julgado de decisão que revoga a gratuidade, a parte deverá efetuar o recolhimento de todas as despesas de cujo (relative pronoun cujo) adiantamento foi dispensada (agentless passive), inclusive as relativas ao recurso interposto, se houver, no prazo fixado pelo juiz, sem prejuízo de aplicação das sanções previstas em lei. (C13.105_16.03.2015)
-
(v)Art. 1º É concedida (agentless passive) isenção de todos os tributos, exclusive a taxa de previdência social (compound noun), que incidam sobre o material, abaixo relacionado (past participle), importado (past participle) para usina hidroelétrica do município de Canápolis, Estado de Minas Gerais, pela Sociedade Brasileira de Eletricidade Siemens Schuckert. Parágrafo único (attributive adjective). O valor (abstract noun) da presente (attributive adjective) importação (abstract noun) é de US$14.000 (quatorze mil dólares). (LO2.000_01.10.1953)
-
(vi)Art. 13. A jurisdição civil (topical adjective) será regida pelas normas processuais brasileiras, ressalvadas (past participle) as disposições específicas (topical adjective) previstas em tratados, convenções ou acordos internacionais de que o Brasil seja parte. (C13.105_16.03.2015)
-
(vii)Art. 1° Fica concedida aos servidores civis e militares do Poder Executivo Federal, da Administração direta (relational adjective), autárquica e fundacional, bem como dos extintos Territórios, a partir de 1° de agosto de 1992, antecipação de reajuste de 20% sobre os vencimentos, soldos e demais retribuições, a ser compensada por ocasião da revisão geral da remuneração dos servidores públicos federais. (LO8.460_17.09.1992)
All these features contribute to the specification, restriction, and detailing of the content conveyed and explain the high informational density and complexity of the syntactic structures found in the LEX-BR-Ius (Ferrari; Marques, 2022).
4.2 Dimension 2: Argumentation
In this dimension, the registers are ranked according to their degree of argumentation (see Figure 3). Those in the positive pole have a higher occurrence of argumentative features, associated with coordinate and dependent clauses, resulting in complex and detailed texts ensuring precision and clarity. Registers in the negative pole, on the other hand, tend to have an absence of these features. Perhaps surprisingly, the register with the highest positive score is Horoscope, and the one with the highest negative score is Websites.
Our corpus and its sub-corpora, except for Códigos, scored in the negative pole. Regardless, every one of our scores is relatively low and close to zero, with values near each other on the dimension scale. By comparing the results obtained with the Legislation register, we observed that the latter is in the negative pole, being more negative than our corpus and, consequently, less argumentative.
These results might indicate that the LEX-BR-Ius (Ferrari; Marques, 2022) is neither decisively argumentative nor completely devoid of argumentation. Thus, the argumentation would not play a primary role in legal texts, as their main objective would be establishing and informing about the rules that govern society. Laws are mandatory, so they must be followed regardless of agreement by the receiver. The argumentation occurs at a stage before the promulgation of the legal text, during the legislative process.
Among the argumentative features, displayed in Table 3, we highlight the occurrence of Infinitive clauses controlled by a preposition (e.g. i), Que clauses controlled by a noun (e.g. i), Que or infinitive clause controlled by a noun (stance) (e.g. ii), Infinitive clause controlled by an adjective (e.g. iii), Que clause controlled by a preposition (e.g. iv), Que clause controlled by adjective (stance) (e.g. v), Infinitive clause controlled by ease or difficulty adjective (e.g. vi). The argumentation features also include Conjunctions: Coordinating (adversative), Articles: Indefinite, Adverbs: Hedge and Comparative, Relative pronouns, and Verbs: Future preterit tense (e.g. vii).
-
(i)Art. 12. Os juízes e os tribunais deverão obedecer à ordem cronológica de conclusão (Nouns: Cognition) para proferir sentença ou acórdão (Infinitive clause controlled by preposition). (C13.105_16.03.2015)
-
(ii)Art. 451. Depois de apresentado o rol de que tratam os §§ 4º e 5º do art. 357, a parte só pode substituir a testemunha: II - que, por enfermidade, não estiver em condições de depor (Que or infinitive clause controlled by noun (stance)) (C13.105_16.03.2015)
-
(iii)Art. 17. Para postular em juízo é necessário ter interesse e legitimidade (Infinitive clause controlled by adjective). (C13.105_16.03.2015)
-
(iv)Art. 26. A cooperação jurídica internacional será regida por tratado de que o Brasil faz parte (Que clause controlled by preposition) e observará: (C13.105_16.03.2015)
-
(v)Art. 95. Cada parte adiantará a remuneração do assistente técnico que houver indicado (Que clause controlled by adjective (stance)), sendo a do perito adiantada pela parte que houver requerido a perícia ou rateada quando a perícia for determinada de ofício ou requerida por ambas as partes. (C13.105_16.03.2015)
-
(vi)Art. 324. O pedido deve ser determinado. § 1º É lícito, porém, formular pedido genérico: II - quando não for possível determinar, desde logo, as consequências do ato ou do fato (Infinitive clause controlled by ease or difficulty adjective) (C13.105_16.03.2015)
-
(vii)Art. 618. Incumbe ao inventariante: II - administrar o espólio, velando-lhe os bens com a mesma diligência que teria (Verbs: Future preterit tense) se seus fossem; (C13.105_16.03.2015)
Due to the difference in our corpus and sub-corpora scores, especially Códigos, we took a closer look at our files. We found some issues in four out of the fifteen features of this dimension. Part of the errors lay within the PALAVRAS parser (Bick, 2001, 2014) that tagged words like meio (environment) and tipo (type) as adverbs when they were nouns. Another mistake of the parser was found in the feature “future preterite verb”: most of those verbs were correctly tagged, however, some adjectives, such as beneficiária (beneficiary) and orçamentária (budgetary), and nouns, like sacaria (set of bags) or maquinaria (machinery), were tagged as verbs.
Some other errors seem to be due to the post-processor. The comparative adverbs feature, for instance, showed many, with only about 10% of the words labeled as such being actual comparative adverbs. The vast majority of the occurrences were actually determiners, such as outro (other/another) and mesmo (the same), and about 6% were adjectives, like maior (bigger) and menor (smaller).
Another issue was found within the feature of infinitive clauses controlled by an adjective. In some cases, the PALAVRAS tagged Roman algorisms correctly as numerals, but the post-processor considered them as adjectives when, in fact, they signal the law’s paragraphs.
As discussed in Section 5, those issues may have biased our results, but it is quite difficult to quantify the degree of this distortion.
4.3 Dimension 3: Involved Versus Informational Production
In Dimension 3, the registers are displayed along the positive pole according to the level of interaction between the participants, and to their informational nature in the negative pole (see Figure 4). The register with the major level of involvement is the Sociolinguistic interview, and the most informational is Emendas à Constituição.
Our corpus, the Leis Complementares, Leis ordinárias and Emendas à Constituição sub-corpora scored in the negative pole. As shown in Table 4, this pole has only one variable: type-token ratio, that reflects lexical richness and diversity. In our corpus, this feature points to a high lexical density reached through a wide range of specialized vocabulary and, complementarily, a lack of interaction. However, to our surprise, some of our sub-corpora, Códigos, Estatutos, and Constituição, are located at the positive pole of this dimension, considerably distant to the others sub-corpora and corpus, having a bigger occurrence of features considered typical of an interactive discourse and, consequently, little lexical diversity. Despite the low score of these sub-corpora in this pole, the results do not fail to raise questions, especially since only they are in this pole and the others are at the extreme of the negative pole. To understand this, we closely checked all occurrences of the features of this dimension. However, we did not find any kind of error in the tagger or in the post-processor. In section 5 we will discuss our hypothesis about those results.
It is important to highlight that some features of the positive pole are indeed recurrent in legal texts. As expected, there are no occurrences of tag questions, contractions, and questions in our corpus. However, there is a significant occurrence of the following features: conclusive coordinate conjunctions (e.g. iv), place adverbs (e.g. v), modal verb have to/have (e.g. vi), pronouns in the third person singular and plural (e.g. iii), and discourse markers (e.g. i and ii).
-
(i)Art. 189. Os atos processuais são públicos, todavia (discourse marker) tramitam em segredo de justiça os processos: (C13.105_16.03.2015)
-
(ii)Art. 242. A citação será pessoal, podendo, no entanto (discourse marker), ser feita na pessoa do representante legal ou do procurador do réu, do executado ou do interessado. (C13.105_16.03.2015)
-
(iii)Art. 231. Salvo disposição em sentido diverso, considera-se dia do começo do prazo: III - a data de ocorrência da citação ou da intimação, quando ela (pronouns: Third person singular, in subject position) se der por ato do escrivão ou do chefe de secretaria; (C13.105_16.03.2015)
-
(iv)Art. 140 - Os empregados contratados há menos de 12 (doze) meses gozarão, na oportunidade, férias proporcionais, iniciando-se, então (Conjunctions: Coordinating (conclusive)), novo período aquisitivo. (C5.452_01.05.1943)
-
(v)Art. 149. As férias poderão ser concedidas, a pedido dos interessados e com aquiescência do armador, parceladamente, nos portos de escala de grande estadia do navio, aos tripulantes ali (place adverbs) residentes. (C5.452_01.05.1943)
-
(vi)Art. 116. O litisconsórcio será unitário quando, pela natureza da relação jurídica, o juiz tiver de (Modals: Ter que/ter de (have to, ought to)) decidir o mérito de modo uniforme para todos os litisconsortes. (C13.105_16.03.2015)
We believe that the significant occurrence of conclusive coordinate conjunctions in our corpus is likely due to the need to fill a gap for other interpretations. The adverbs of place, by their turn, would exercise the function of locating and specifying the subject/object of the norm. As for the modal verbs, they dictate the rule/ the expected behavior of the subject to whom the norm is addressed.
4.4 Dimension 4: Directive Discourse
In dimension 4, the registers are characterized by the presence of guidelines/orders/instructions for executing tasks. As displayed in Figure 5, the most prototypical register of this dimension is Recipes, while the LEX-BR-Ius corpus and its sub-corpora scored at the negative pole, being, thus, marked by the absence of directive discourse.
The features of this dimension are displayed in Table 5 and include imperative and present subjunctive mood used in a directive sense (e.g. i, ii), facilitation verbs (e.g. iii), concrete nouns (e.g. iii), and coordinating clausal conjunctions (e.g. iv).
-
(i)Art. 4º. Acrescente-se (Verbs: Imperative mood) o § 3º. ao art. 142 da Constituição: (EC18_05.02.1998)
-
(ii)Art. 10. O juiz não pode decidir, em grau algum de jurisdição, com base em fundamento a respeito do qual não se tenha (Verbs: Present subjunctive mood) dado às partes oportunidade de se manifestar, ainda que se trate de matéria sobre a qual deva (Verbs: Present subjunctive mood) decidir de ofício. (C13.105_16.03.2015)
-
(iii)Art. 1º É concedida (Verbs: Facilitation) isenção de todos os tributos, exclusive a taxa de previdência social, que incidam sôbre o material, abaixo relacionado, importado para usina (Nouns: Concrete) hidroelétrica do município de Canápolis, Estado de Minas Gerais, pela Sociedade Brasileira de Eletricidade Siemens Schuckert. a) parte hidráulica completa; b) parte elétrica: gerador (Nouns: Concrete) e quadro (Nouns: Concrete). (LO2.000_01.10.1953)
-
(iv)Art. 1º Os arts. 2º , 3º , 6º , 12, 13, 14, 15, 16, 17, 20, 21, 22, 23, 24, 28 e 36 da Lei nº 8.742, de 7 de dezembro de 1993, passam a vigorar com a seguinte redação: “Art. 3º Consideram-se entidades e organizações de assistência social aquelas sem fins lucrativos que, isolada ou (Conjunctions: Coordinating (clausal)) cumulativamente, prestam atendimento e assessoramento aos beneficiários abrangidos por esta Lei, bem como as que atuam na defesa e garantia de direitos. […]” (LO12.435_06.07.2011).
The placement of our corpus and sub-corpora in the dimension scale was quite intriguing, so we inspected each feature occurrence to check the post-processor outcomes and the PoS tag accuracy.
First, for some unknown reason, the feature subject omission was not counted in our corpus. A minor issue was also found in the tagging of the subjunctive present: about 0.6% of the more than 18 thousand occurrences were, actually, indicative futures.
Conversely, the features of imperative and subjunctive mood exhibited several errors. The post-processor counted a total of 345 tokens as imperative as well as subjunctive, about 98% of which were wrongly tagged. Some proper nouns, like America and Alegre, for example, were mistakenly labeled as imperative and subjunctive. Additionally, some acronyms, such as “CIE” (Carteira de Identificação Estudantil), “CAPES” (Coordenação de Aperfeiçoamento de Pessoal de Nível Superior), and “Fies” (Fundo de Financiamento Estudantil), as well as nouns like ente (entity), repasse (transfer) and sede (headquarters), were incorrectly tagged as verbs. Lastly, some subordinate subjunctives were mistakenly classified as imperatives.
We identified orthographic errors in one of our texts (LO62_05.06.1935) that probably contributed to the incurrence of the tagger output in this specific text. Apart from this text, in which the tagging errors are justified and expected due to the orthographic issues found, the vast majority of the before-mentioned errors befall the PALAVRAS tagger (Bick, 2000, 2014) wrongly PoS classification in our corpus.
As stated before, the incorrect tags could have influenced the results of the MD analysis, but it is quite difficult to specify how it impacted our results.
4.5 Dimension 5: Future Versus Past Time Orientation
Dimension 5 describes the registers in a scale ranging from future-oriented registers (positive pole) to past-oriented ones (negative pole). In this dimension, our Códigos sub-corpus scored as the most prototypical of the positive pole, while the most prototypical of the negative pole is Youth fiction, as shown in Figure 6.
Among this dimension’s features, displayed in Table 6, we highlight the strong presence of Verbs: Future present tense (e.g. i, ii), Future subjunctive mood (e.g. iii), and Subordinating (conditional) clause (e.g. ii) in our corpus. Other features include Conjunctions: Coordinating (ou) and Coordinating (phrasal), use of modals (duty and power), and Likelihood adverbs.
-
(i)Art. 12. Os juízes e os tribunais atenderão (future present tense), preferencialmente, à ordem cronológica de conclusão para proferir sentença ou (coordinating ou) acórdão. (C13.105_16.03.2015)
-
(ii)Art. 39. O pedido passivo de cooperação jurídica internacional será (future present tense) recusado se (Subordinating (conditional) clause) configurar manifesta ofensa à ordem pública. (C13.105_16.03.2015)
-
(iii)Art. 21. Compete à autoridade judiciária brasileira processar e julgar as ações em que: II - no Brasil tiver (Future subjunctive mood) de ser cumprida a obrigação; (C13.105_16.03.2015)
Such features establish how the laws must be applied and their consequences in concrete cases after promulgation. That said, it is interesting to notice a big difference between our corpus score and that of the Legislation register, which is drastically lower.
It is important to highlight that we identified several errors in the feature Adverb: Likelihood. The issue resides in the post-processor as the parser correctly tagged the PoS of the 10,896 occurrences of this feature. However, out of these, only 544 were adverbs, like aproximadamente (approximately). The post-processor misclassified the remaining 10,352 occurrences as adverbs of likelihood when they were actually prepositions, e.g. de (of), nouns, e.g. tipo (type), and verbs in the indicative future. In the latter, the third person singular of the verb ser (to be), será, accounted for about 93% (10,174 cases) of these occurrences. This probably significantly impacted the analysis results, but we can only speculate about the extent of these incorrect labels on our findings.
4.6 Dimension 6: Reported Discourse
Finally, the last dimension of variation ranks the registers according to their degree of reported discourse. The register that scored most positively in this dimension is Church liturgy, and the most negative one is Recipes.
In this dimension, our texts scored in both poles. While Códigos is located at the positive pole of this dimension, our corpus, and its other sub-corpora, are spread along the negative pole. Based on these results we concluded that indirect discourse is not a striking characteristic of the corpus, and direct discourse is generally preferred. By their very nature, legal texts are made to be as clear as possible, and do not report what has been said in other texts, as they are the ones who dictate the rules to their receiver. The features of this dimension are displayed in Table 7 below.
We also identified errors in the parser and post-processor in this dimension. In spite of having issues with the tagging and counting, most of the subordinating final clauses (e.g. iii) and the modals Haver que/Haver de (e.g. ii) didn’t show issues. As for the feature Pronouns: Rare in object position (e.g. ii), that comprises object pronouns such as datives lhe, lhes, and accusatives lo and la, the absolute majority of the pronouns were correctly tagged, except for two occurrences of the conjunction si, outdated form of se (if), in the LO26_30.12.1891.
Regarding the occurrences of the feature public verbs (e.g. i), they included verbs such as requerer (require), postular (postulate), revogar (revogate), and publicar (publish). Out of the more than 29,000 occurrences of this feature, only a very little percentage presented errors. The mistakes laid with some adjectives, as pública (public), and nouns, such as república (republic), that were misjudged as verbs. In these cases, we conjecture that this mistake was probably due to the lack of the acute accent in some parts of our texts, what might have led the tagger to identify, for example, the noun pública (public) as the verb publica (to publish) texts.
The most puzzling feature was Verbs in the second person. In theory, laws should not present verbs in the second person, but, oddly, the post-processor counted 1,360 occurrences of this feature in our corpus. Upon checking the files, we constated that all the occurrences were incorrectly tagged. Nouns, such as custas (costs), semoventes (movable property able to move on their own, e.g. animals), and pertences (belongings), as well as acronyms, like “CAPES” and “Fies”, and adjectives, as tributáveis (taxable), aplicáveis (applicable), and federais (federals) were mistakenly attributed to second-person verb tags. We aren’t sure if a small number of wrong occurrences as those 1,360 would bias the analysis or their effect on our corpus and sub-corpora scores in this dimension.
-
(i)Art. 5o Revoga-se (Verbs: Public) a Lei Complementar no 92, de 23 de dezembro de 1997. (LC99_20.12.1999)
-
(ii)Art. 687. A habilitação ocorre quando, por falecimento de qualquer das partes, os interessados houverem de (Modals: Haver que/haver de (have to, ought to)) suceder-lhe (Pronouns: Rare in object position) no processo. (C13.105_16.03.2015)
-
(iii)Art. 326. É lícito formular mais de um pedido em ordem subsidiária, a fim de que (Subordinating (final) clause) o juiz conheça do posterior, quando não acolher o anterior. (C13.105_16.03.2015)
4.8 Significance of the Variation
The ANOVA tests for the sub-corpora of LEX-BR-Ius - in a version that excluded the Constituição register since it has only a single file in this category - are exhibited in Table 8.
ANOVAs of LEX-BR-Ius sub-corpora for CBVR dimensions of variation (Berber Sardinha; Kauffmann; Acunzo, 2014) (N = 753, df = 4)
Four out of the six BP dimensions (1, 3, 4, and 5) featured significant differences among our sub-corpora, although with discreet values of R².
The ANOVA performed with a sample of our corpus and the CBVR registers is shown in Table 9.
ANOVAs of a sample of LEX-BR-Ius corpus with CBVR registers and dimensions of variation (Berber Sardinha; Kauffmann; Acunzo, 2014) (N = 980, df = 48)
The ANOVA results indicate statistically significant differences between our corpus and the CBVR registers in all BP dimensions with high values of R², capturing between 50.5% and 78.5% of the variation.
5 Discussion
As shown in Section 4, some of our results were rather unexpected and puzzling. Among other reasons, we believe that the characteristics of the LEX-BR-Ius corpus (Ferrari; Marques, 2022), such as the number of words per text and per sub-corpora, contributed to these results. Even though we normalized our data and calculated the Z-scores to avoid skewing the analysis, our corpus design might have played a role in the discrepancies found. This may justify the results of the first dimension, Oral vs. Literate discourse. While Leis Ordinárias and Emendas à Constituição have the smallest texts, Códigos have the biggest ones. This reasoning also applies to dimension 5, Future vs. Past orientation, where Emendas à Constituição has a relatively lower score when compared to the rest of the corpus. Additionally, the different compilation criteria and composition of our corpus and the CBVR (Berber Sardinha; Kauffmann; Acunzo, 2014) might have interfered with the results and help explain the differences in factor loading between our corpus and sub-corpora and the Legislation register.
Another reason for these unexpected results are the PALAVRAS (Bick, 2000, 2014) errors, detected in the second, fourth, and sixth dimensions, and of the post-processor PT Tag Count (Berber Sardinha, 2013a), detected in dimensions two and five. Despite PALAVRAS being a state-of-the-art tagger, there were several errors in the PoS tagging in our corpus. These errors illustrate that legal language is a specific variety, requiring additional adjustments in the software to capture its nuances. On the other hand, the post-processor, now in its 4th version, has been fine-tuned in the last few years, and we hope additional corrections with legal language in mind would help improve the MD description of this register. As said, it’s difficult to estimate the range of the bias those errors generated in the results.
Nevertheless, the most intriguing findings are those of dimension 3, Involved vs. Informational production. Although no error was identified in the parser nor the post-processor, it is quite odd that a part of the corpus is placed at the extremity of the negative pole, indicating strong informational content, while another part is at the beginning of the positive pole, denoting some involvement, and our situational analysis shows that the involvement of the participants in the legal text is practically null.
These findings deserve a separate reflection as legal texts are not expected to figure as involved productions. Such incongruences may be explained by factors inherent to the corpus: its design (1) and the situational features of the types of law compressed by it (2); the MD approach requirements (3); and the feature discourse marker label (4).
1 The corpus design: the LEX-BR-Ius (Ferrari; Marques, 2022) is undoubtedly unbalanced. As shown in Section 3.1, the number of texts and words per sub-corpora greatly diverges. The sub-corpora with the largest number of words per text in our corpus, namely Códigos, Constituição, and Estatutos, comprise 28% of the total words in the corpus (934,390) in 25 texts, standing out considerably from the other sub-corpora in this regard. This fact may have biased the sample towards some features, which would justify disparate results between the entire corpus set and its internal sections.
2 Sub-corpora’s situational characteristics: the Constituição, the Códigos, and the Estatutos sub-corpora present, in general, the longest texts in our corpus. This is due to the function of these texts:
- The Constituição establishes the basis of Brazilian society. It is the main law that regulates the state and the division of powers, establishes the fundamental rights, and lays the rules of the Brazilian legal system.
- The Códigos establish the norms and principles that guide broad areas of law, such as criminal law and tax law.
- The Estatutos regulate specific groups of the population, e.g. the elderly.
The other laws (Emendas à Constituição, Leis Complementares, and Leis Ordinárias) are subject to the rules established by these three law types. While the first only complements, adds, or makes occasional modifications in specific parts of the Constitutional text, the second regulates and elaborates on issues expressly established in the Constitution as their object, and the latter regulates everything that is not covered by the other types of law. To fulfill its objectives, these norms must be extremely explanatory, detailed, and comprehensive, aiming to cover as many scenarios as possible and accurately determine how to act in them, what justifies their size.
3 MD requirements: to perform an MD, the number of texts and words per register in the corpus under analysis should be similar. Additionally, in the additive MD, the corpus under analysis must be compiled following the same criteria and using the same software as the base study corpus. Despite using the same software as the base study, our corpus is 1 register with 755 texts and about 3.3 million words. The CBVR (Berber Sardinha; Kauffmann; Acunzo, 2014) covers 48 registers, with 20 texts per register, totaling 960 texts and 5.6 million words. Due to that, to compare the corpora, we normalized and standardized the data to prevent any bias in the results arising from discrepancies in numbers (refer to Section 3). Despite these efforts, we cannot disregard this factor entirely.
4 Discourse marker feature: we clarify that the feature discourse marker refers to connectives, which, by the very nature of the legal text, are expected to occur. Among them, we highlight the high occurrence of todavia, portanto, and entretanto in our corpus. These connectives are typical of formal and informational written discourse, used to connect coordinate and subordinate clauses, reflecting a high degree of grammatical complexity in legal texts. For this reason, we believe they do not necessarily mark an interactive “involvement” but a level of argumentation that explains and describes the norms in greater detail.
6 Conclusions
In this paper, we presented an additive MD analysis of the LEX-BR-Ius corpus (Ferrari; Marques, 2022), a representative corpus of Brazilian federal statutory laws, its results, and issues. To perform the analysis, we added our corpus to the CBVR (Berber Sardinha; Kauffmann; Acunzo, 2014) in all Brazilian Portuguese dimensions of variation identified by Berber Sardinha, Kauffmann, and Acunzo (2014).
In all dimensions, our corpus computed unique factor loadings and statistically significant differences when compared to the CBVR registers, confirming Brazilian laws as a register. As for our sub-corpora, they exhibited different factor loadings in all dimensions. While this difference is discrete in some dimensions, there are some remarkable differences in others, being as far as located in different poles in dimensions 2, 3, and 6. The statistical tests performed in our sub-corpora showed little but significant differences between them in dimensions 1, 3, 4, and 5. Due to that, despite their different situational characteristics and factor loadings, we cannot affirm that they are sub-registers.
Among the challenges we faced were our corpus’s lack of balance and its unique textual and situational traits. We also pointed out some methodological issues of the multidimensional approach, and both the PALAVRAS (Bick, 2000, 2014) and the PT Tag Count (Berber Sardinha, 2013a) errors, which might explain some of the oddities in our findings. In navigating these challenges with a highly specialized corpus, this study also established certain “yardsticks” to adjust the original settings of these tools for the analysis of legal discourse.
Despite these issues, this study demonstrates the effectiveness of additive MD analysis in tracing linguistic patterns in BP’s legal language. However, further adjustments in the parser and post-processor are necessary to enhance accuracy in future research.
Data Availability Statement
The data necessary to reproduce the findings of this study, specifically the data from the LEX-BR-Ius, is available for download at http://www.letras.ufmg.br/profs/luciaferrari/lex-br-ius. This study also used third-party data from the CBVR Multidimensional Analysis. Part of this data can be found in the work of Berber Sardinha, Kauffmann, and Acunzo (2014), while the remainder is subject to licensing restrictions and cannot be publicly shared by the authors.
References
- ATKINS, S.; CLEAR, J. Corpus Design Criteria. Literary and Linguistic Computing, v. 7, n. 1, p. 1-16, 1992.
- ATKINSON, D. Scientific Discourse Across History: A Combined Multidimensional/ Rhetorical Analysis of the Philosophical Transactions of the Royal Society of London. In: CONRAD, S.; BIBER, D. (org.). Variation in English: Multi-Dimensional Studies. Harlow: Longman, 2001. p. 45-65.
- ATKINSON, D. The Evolution of Medical Research Writing from 1735 to 1985: The Case of the ‘Edinburgh Medical Journal’. Applied Linguistics, v. 13, p. 337-74, 1992.
- ATKINSON, D. The Philosophical Transactions of the Royal Society of London, 1675-1975: A Sociohistorical Discourse Analysis. Language in Society, v. 25, p. 333-71, 1996.
- BARBERA, M.; ONESTI, C. Scheda progetto di ricerca n. 9. Corpus Jus Jurium. In: DIADORI, P. Progetto JURA: la formazione dei docenti di lingua e traduzione in ambito giuridico italo. Perugia: Guerra Edizioni, 2009. p. 349-351.
-
BERBER SARDINHA, T. A abordagem metodológica da Análise Multidimensional. Gragoatá, Niterói, n. 29, p. 107-125, 2. sem. 2010. Available at: Available at: https://periodicos.uff.br/gragoata/article/view/33077 Accessed on: 19 Jan. 2022.
» https://periodicos.uff.br/gragoata/article/view/33077 - BERBER SARDINHA, T. A Text Typology of Social Media. Register Studies, v. 4, p. 138-170, 2022.
- BERBER SARDINHA, T. AI-Generated vs Human-Authored Texts: A Multidimensional Comparison. Applied Corpus Linguistics, v. 4, e100083, 2024.
- BERBER SARDINHA, T. et al Adding Registers to a Previous Multi-Dimensional Analysis. In: BERBER SARDINHA, T.; VEIRANO PINTO, M. (ed.). Multi-Dimensional Analysis: Research Methods and Current Issues. London: Bloomsbury, 2019. p. 165-186.
- BERBER SARDINHA, T. Informatividade, interatividade e narratividade na reunião de negócios - Análise Multidimensional e palavras-chave. DIRECT Papers, 52. São Paulo, SP: LAEL/ PUC-SP; Liverpool: ELSU/ University of Liverpool, 2003.
- BERBER SARDINHA, T. Pós-processador PT Tag Count (Version 4) [Computer Software]. 2013a.
- BERBER SARDINHA, T. Variação entre registros da internet. In: SHEPHERD, T. G.; SALIÉS, T. G. (ed.). Linguística da Internet São Paulo: Contexto, 2013b. p. 55-85.
-
BERBER SARDINHA, T.; ALENCAR, A. L. S.; SILVA, C. S.; GIL, C. B.; LOPES, M. J. F.; HUGHES, S. A. S. #eunaovoutomarvacina: uma abordagem da linguística de corpus e da análise multimodal imagética. Intercâmbio, v. 51, e58515, 2022. Available at: Available at: https://revistas.pucsp.br/index.php/intercambio/article/view/58515 Accessed on: 9 June 2024.
» https://revistas.pucsp.br/index.php/intercambio/article/view/58515 - BERBER SARDINHA, T.; KAUFFMANN, C. H.; ACUNZO, C. M. Dimensions of Register Variation in Brazilian Portuguese. In: VEIRANO PINTO, M. (ed.). Multi-Dimensional Analysis, 25 Years On: A Tribute to Douglas Biber. Amsterdam; Philadelphia: John Benjamins Publishing Company, 2014. p. 35-80.
- BÉRTOLI-DUTRA, P. Linguagem da música popular anglo-americana de 1940-2009 2010. Tese (Doutorado em Linguística Aplicada e Estudos da Linguagem) - Pontifícia Universidade Católica de São Paulo, São Paulo, 2010.
- BIBER, D. A Comparison of British and American Writing. American Speech, v. 62, p. 99-119, 1987.
- BIBER, D. Dimensions of Register Variation: A Cross-Linguistic Perspective. Cambridge: CUP, 1995.
- BIBER, D. et al Speaking and Writing in the University: A Multi-Dimensional Comparison. TESOL Quarterly, v. 36, n. 1, p. 9-48, 2002.
- BIBER, D. Representativeness in Corpus Design. Literary and Linguistic Computing, Oxford, v. 8, n. 4, p. 243-257, 1993.
- BIBER, D. Variations Across Speech and Writing Cambridge: CUP, 1988.
- BIBER, D; CONRAD, S. Register, Genre, and Style Cambridge: CUP, 2009.
- BIBER, D.; FINEGAN, E. Diachronic Relations Among Speech-Based and Written Registers in English. In: CONRAD, S. M.; BIBER, D. (ed.). Variation in English: Multi-Dimensional Studies. Harlow: Longman, 2001. p. 66-83.
- BIBER, D.; FINEGAN, E. Diachronic Relations Among Speech-Based and Written Registers in English. In: NEVALAINEN, T.; KAHLASTARKKA, L. (ed.). To Explain the Present: Studies in the Changing English Language in Honour of Matti Rissanen. (Mémoires de la Société Néophilologique de Helsinki 52). Helsinki: Société Néophilologique de Helsinki, 1997. p. 253-275.
- BIBER, D.; FINEGAN, E. Drift and the Evolution of English Style: A History of Three Genres. Language, v. 65, p. 487-517, 1989.
- BIBER, D.; FINEGAN, E. Drift in Three English Genres from the 18th to the 20th Centuries: A Multidimensional Approach. In: KYTÖ, M.; IHALAINEN, O.; RISSANEN, M. (ed.). Corpus Linguistics, Hard and Soft: Proceedings of the Eighth International Conference on English Language Research on Computerized Corpora. Amsterdam: Rodopi, 1988. p. 83-101.
- BIBER, D.; FINEGAN, E. The Linguistic Evolution of Five Written and Speech-Based English Genres from the 17th to the 20th Centuries. In: RISSANEN, M.; IHALAINEN, O.; NEVALAINEN, T.; TAAVITSAINEN, I. (ed.). History of Englishes: New Methods of Interpretations in Historical Linguistics. Berlin; New York: Mouton de Gruyter, 1992. p. 688-704.
- BIBER, D.; HARED, M. Dimensions of Register Variation in Somali. Language Variation and Change, v. 4, n. 1, p. 41-75, 1992a.
- BIBER, D.; HARED, M. Literacy in Somali: Linguistic Consequences. Annual Review of Applied Linguistics, v. 12, p. 260-82, 1992b.
- BICK, E.; PALAVRAS, A Constraint Grammar-Based Parsing System for Portuguese. In: BERBER SARDINHA, T.; FERREIRA, T. (ed.). Working with Portuguese Corpora London: Bloomsbury, 2014. p. 279-302.
- BICK, E. The Parsing System “Palavras”: Automatic Grammatical Analysis of Portuguese in a Constraint Grammar Framework. 2000. Thesis (Doctor of Philosophy) - Aarhus University, Denmark, 2000.
- BRAZ, A. A. B. Representações contemporâneas da sustentabilidade: uma análise multidimensional lexical discursiva como contribuição para o Portal Multimodal/Multilíngue para Avanço da Ciência Aberta nas Humanidades. 2023. Dissertação (Mestrado em Linguística Aplicada e Estudos da Linguagem) - Pontifícia Universidade Católica de São Paulo, São Paulo, 2023.
- BREZINA, V. Statistics in Corpus Linguistics: A Practical Guide. Cambridge: Cambridge University Press, 2018.
- CAO, Y.; XIAO, R. A Multi-Dimensional Contrastive Study of English Abstracts by Native and Non-Native Writers. Corpora, v. 8, n. 2, p. 209-234, 2013.
-
CHOVANEC, J. Grammar in the Law. In: CHAPELLE, C. A. (ed.). The Encyclopedia of Applied Linguistics New Jersey: Blackwell Publishing Ltd, 2013. p. 1-8. Available at: Available at: https://www.academia.edu/3389303/Grammar_in_the_Law Accessed on: 6 Sept. 2020.
» https://www.academia.edu/3389303/Grammar_in_the_Law - CONDE, H. M. D. A. Escolhas léxico-gramaticais em composições de alunos avançados de inglês originários de instituições de ensino bilíngües e monolíngües - um estudo multidimensional baseado em corpus. 2002. Dissertação (Mestrado em Linguística Aplicada e Estudos da Linguagem) - Pontifícia Universidade Católica de São Paulo, São Paulo, 2002.
- CONNOR, U.; UPTON, T. A. Linguistic Dimensions of Direct Mail Letters. In: LEISTYNA, P.; MEYER, C. F. (org.). Corpus Analysis: Language Structure and Language Use. Amsterdam; New York: Rodopi, 2003. p. 71-86.
- CONRAD, S. Investigating Academic Texts with Corpus-Based Techniques: An Example from Biology. Linguistics and Education, v. 8, n. 299-326, 1996.
- CONRAD, S. M. Expanding Multi-Dimensional Analysis with Qualitative Research Techniques. In: BERBER SARDINHA, T.; VEIRANO PINTO, M. (ed.). Multi-Dimensional Analysis, 25 Years On: A Tribute to Douglas Biber. Amsterdam; Philadelphia, PA: John Benjamins, 2014. p. 344-411.
- CONRAD, S. M. Variation Among Disciplinary Texts: A Comparison of Textbooks and Journal Articles in Biology and History. In: CONRAD, S. M.; BIBER, D. (ed.). Variation in English: Multi-Dimensional Studies. Harlow: Longman, 2001. p. 94-107.
-
DAHLMAN, R. C. Specialità del linguaggio giuridico italiano. Stockholm: Stockholm University, 2006. Available at: Available at: https://www.researchgate.net/publication/27818945_Specialita_del_linguaggio_giuridico_italiano Accessed on: 6 Sept. 2020.
» https://www.researchgate.net/publication/27818945_Specialita_del_linguaggio_giuridico_italiano -
DELFINO, M. C. N. Análise Multidimensional: os números na linguística. Cadernos de Linguística, v. 2, n. 4, p. 1-21, 2021. Available at: Available at: https://www.researchgate.net/publication/354864891_Analise_Multidimensional_os_numeros_na_Linguistica Accessed on: 19 Jan. 2022.
» https://www.researchgate.net/publication/354864891_Analise_Multidimensional_os_numeros_na_Linguistica - DELFINO, M. C. N.; BERBER SARDINHA, T.; COLLENTINE, J. Dimensões de variação lexical e acústica na música popular em inglês: um estudo baseado em corpus. Cadernos de Estudos Linguísticos (UNICAMP), v. 65, e23025, 2023.
- ESCARABELIN, L. F. A linguagem verbal de videogames em uma perspectiva multidimensional e da ciência aberta 2022. Dissertação (Mestrado em Lingüística Aplicada e Estudos da Linguagem) - Pontifícia Universidade Católica de São Paulo, São Paulo, 2022.
-
FERRARI, L. A.; CUNHA, E. L. T. P. Reflexões metodológicas sobre datasets e linguística de corpus: uma análise preliminar de dados legislativos. Domínios de Lingu@gem, v. 16, n. 4, p. 1571-1607, 2022. Available at: Available at: https://seer.ufu.br/index.php/dominiosdelinguagem/article/view/64146 Accessed on: 23 Sept. 2022.
» https://seer.ufu.br/index.php/dominiosdelinguagem/article/view/64146 -
FERRARI, L. A.; MARQUES, C. G. F. O LEX-BR-Ius: arquitetura e decisões na compilação de um corpus representativo das leis federais brasileiras. ANTARES, v. 14, n. 34, 2022. Available at: Available at: http://www.ucs.br/etc/revistas/index.php/antares/article/view/11150/5328 Accessed on: 19 Dec. 2022.
» http://www.ucs.br/etc/revistas/index.php/antares/article/view/11150/5328 - FORCHINI, P. Movie Language Revisited: Evidence from Multi-Dimensional Analysis and Corpora. Bern: Peter Lang, 2012.
- GEISLER, C. Gender-Based Variation in Formal Spoken American English. In: GRAMMATIK I FOKUS, 2001, Lund. ASLA SYMPOSIUM, 2001, Uppsala. Proceedings […]. [S. l.]: ASLA, 2001a.
- GEISLER, C. Gender-Based Variation in Nineteenth-Century Letterwriting. In: NORTH AMERICAN SYMPOSIUM ON CORPUS LINGUISTICS AND LANGUAGE TEACHING, 3., 2001, Boston. Proceedings […]. Boston: Brill Academic Pub, 2001b.
- GEISLER, C. Investigating Register Variation in Nineteenth-Century English: A Multi-Dimensional Comparison. In: NORTH AMERICAN SYMPOSIUM ON CORPUS LINGUISTICS AND LANGUAGE TEACHING, 2., 2000, Flagstaff. ICAME, 21., 2000, Sydney. Proceedings […]. Boston: Sciendo, 2000.
- GOŹDŹ-ROSZKOWSKI, S. Legal Language. In: CHAPELLE, C. A. (org.). The Encyclopedia of Applied Linguistics [S. l.]: John Wiley e Sons, 2012. p. 3281-3287.
- GOŹDŹ-ROSZKOWSKI, S.; PONTRANDOLFO, G. Facing the Facts: Evaluative Patterns in English and Italian Judicial Language. In: BHATIA, V.; GARZONE, G.; SALVI, R. (ed.). Language and Law in Professional Discourse Newcastle upon Tyne: Cambridge Scholars, 2014. p. 10-28.
-
HARDIE, A. Modest XML for Corpora: Not a Standard, But a Suggestion. ICAME Journal, v. 38, n. 1, p. 73-103, 2014. Available at: https://doi.org/10.2478/icame-2014-0004. Accessed on: 20 July 2021.
» https://doi.org/10.2478/icame-2014-0004 - HELT, M. A Multi-Dimensional Comparison of British and American Spoken English. In: CONRAD, S.; BIBER, D. (org.). Variation in English: Multi-Dimensional Studies. Harlow: Longman, 2001. p. 171-184.
-
HO, D. Notepad++ (Version 7.8.9) [Computer Software]. 2020. Available at: Available at: https://notepad-plus-plus.org/downloads/v7.8.9/ Accessed on: 5 Mar. 2020.
» https://notepad-plus-plus.org/downloads/v7.8.9/ - JONSSON, E. Conversational Writing: A Multi-Dimensional Study of Synchronous and Supersynchronous Computer-Mediated Communication. Frankfurt am Main; New York: Peter Lang, 2016.
- KAUFFMANN, C. H. Linguística de corpus e estilo: análises multidimensional e canônica na ficção de Machado de Assis. 2020. Tese (Doutorado em Linguística Aplicada e Estudos da Linguagem) - Pontifícia Universidade Católica de São Paulo, São Paulo, 2020.
- KAUFFMANN, C. H. O corpus do jornal: variação lingüística, gêneros e dimensões da imprensa diária escrita. 2005. Dissertação (Mestrado em Linguística Aplicada e Estudos da Linguagem) - Pontifícia Universidade Católica de São Paulo, São Paulo, 2005.
- KIM, Y.-J.; BIBER, D. A Corpus-Based Analysis of Register Variation in Korean. In: BIBER, D.; FINEGAN, E. (org.). Sociolinguistic Perspectives on Register Oxford: Oxford University Press, 1994. p. 157-181.
- LAMB, W. Scottish Gaelic Speech and Writing: Register Variation in an Endangered Language. Belfast: Cló Ollscoil na Banríona, 2008.
- LAMB, W. Speech and Writing in Scottish Gaelic: A Study of Register Variation in an Endangered Language. 2002. Thesis (Doctorate in Linguistics) - University of Edinburgh, Edinburgh, 2002.
- LEECH, G. New Resources, or Just Better Old Ones? The Holy Grail of Representativeness. In: HUNDT, M.; NESSELHAUF, N.; BIEWER, C. (ed.). Corpus Linguistics and the Web Amsterdam: Brill Rodopi, 2007. p. 133-50.
- LORZ, R. A. Creating Law with Language - Crossing Borders and Connecting Disciplines from the Perspective of Legislative Practice. In: VOGEL, F. (ed). Legal Linguistics Beyond Borders: Language and Law in a World of Media, Globalisation and Social Conflict Relaunching the International Language and Law Association (ILLA). Berlin: Duncker e Humblot GmbH, 2019. p. 5-8.
- MCENERY, T.; HARDIE, A. Corpus Linguistics: Method, Theory, and Practice. Cambridge: Cambridge University Press, 2012.
-
MCENERY, T.; XIAO, R.; TONO, Y. Corpus-Based Language Studies: An Advanced Resource Book. London: Routledge, 2006. Available at: Available at: https://www.lancaster.ac.uk/fass/projects/corpus/ZJU/xCBLS/CBLS.htm Accessed on: 5 Mar. 2020.
» https://www.lancaster.ac.uk/fass/projects/corpus/ZJU/xCBLS/CBLS.htm - OLIVEIRA, L. P. Variação intercultural na escrita: contrastes multidimensionais em inglês e português. 1997. Tese (Doutorado em Linguística Aplicada) - Pontifícia Universidade Católica de São Paulo, São Paulo, 1997.
- ONESTI, C. Methodology for Building a Text-Structure Oriented Legal Corpus. Comparative Legilinguistics, v. 8, p. 37-48, 2011.
- PARODI, G. Variation Across Registers in Spanish: Exploring the El-Grial PUCV Corpus. In: PARODI, G. (org.). Working with Spanish Corpora London: Continuum, 2007. p. 11-53.
- REY, J. Changing Gender Roles in Popular Culture: Dialogue in Star Trek Episodes from 1966 to 1993. In: CONRAD, S.; BIBER, D. (org.). Variation in English: Multi-Dimensional Studies. Harlow: Longman, 2001. p. 138-156.
-
RICHARD, I. Is Legal Lexis a Characteristic of Legal Language? Lexis, n. 11, 2018. Available at: Available at: https://www.researchgate.net/publication/324949333_Is_legal_lexis_a_characteristic_of_legal_language Accessed on: 6 Sept. 2020.
» https://www.researchgate.net/publication/324949333_Is_legal_lexis_a_characteristic_of_legal_language - ROSSINI FAVRETTI, R.; TAMBURINI, F.; MARTELLI, E. Words from Bononia Legal Corpus. International Journal of Corpus Linguistics, v. 6 (Special Issue), p. 13-34, 2001.
- SANTOS, V. B. M. P. D. As características léxico-gramaticais de um manual de gestão a partir da análise multidimensional de Biber. The ESPecialist, v. 24, n. 2, p. 201-227, 2003.
-
SAS INSTITUTE. SAS® OnDemand for Academics [Computer Software]. 2022. Available at: Available at: https://welcome.oda.sas.com/ Accessed on: 5 Mar. 2023.
» https://welcome.oda.sas.com/ - SILVEIRA, M. S. D. Contrastes interculturais na linguagem de negócios: Análise Multidimensional em L1 e L2. Projeto de Iniciação Científica. Rio de Janeiro: PUC-Rio, 1997.
- SINCLAIR, J. Trust the Text: Language, Corpus, and Discourse. London: Routledge, 2004.
- SOUZA, R. C. Dimensions of Variation in TIME Magazine. In: BERBER SARDINHA, T.; VEIRANO PINTO, M. (ed.). Multi-Dimensional Analysis, 25 Years On: A Tribute to Douglas Biber. Amsterdam; Philadelphia, PA: John Benjamins, 2014. p. 177-193.
-
SVOBODOVÁ, I. Modalidade não epistêmica na linguagem jurídica: um estudo contrastivo. Caligrama, Belo Horizonte, v. 22, n. 2, p. 103-133, 2017. Available at: http://dx.doi.org/10.17851/2238-3824.22.2.103-133. Accessed on: 6 Sept. 2020.
» https://doi.org/10.17851/2238-3824.22.2.103-133 - TIERSMA, P. Legal Language Chicago: The University of Chicago Press, 1999.
- VEIGA, A. T. As dimensões da fé: sete religiões mundiais em análise multidimensional lexical. 2021. Tese (Doutorado em Lingüística Aplicada e Estudos da Linguagem) - Pontifícia Universidade Católica de São Paulo, São Paulo, 2021.
- VEIRANO PINTO, M. Dimensions of Variation in North American Movies. In: BERBER SARDINHA, T.; VEIRANO PINTO, M. (ed.). Multi-Dimensional Analysis, 25 Years On: A Tribute to Douglas Biber. Amsterdam; Philadelphia, PA: John Benjamins, 2014. p. 109-147.
- ZUPPARDI, M. C. Collocation Dimensions in Academic English 2020. Thesis (Doctorate in Applied Linguistics and Language Studies) - Pontifícia Universidade Católica de São Paulo, São Paulo, 2021.
-
Funding:
This research was partially financed by: CAPES (grant nº 88887.626989/2021-00) and FAPEMIG (grants PROBIC and PIC-JR-FAPEMIG).
-
Use of AI:
In this scientific work, generative artificial intelligence (AI) has not been used.
-
Reviews:
As part of the commitment made by the Brazilian Journal of Applied Linguistics to Open Science, the journal publishes the reviews issued regarding its published works, when authorized by all parties involved.
Edited by
-
Responsible Editor:
Andréa Machado de Almeida Mattos, Universidade Federal de Minas Gerais (UFMG), Belo Horizonte, Minas Gerais/MG, Brasil. Lattes: http://lattes.cnpq.br/7749222257907067, ORCID: https://orcid.org/0000-0003-3190-7329, e-mail: andreamattos@ufmg.br.








Source: The authors (2024).
Source: The authors (2024).
Source: The authors (2024).
Source: The authors (2024).
Source: The authors (2024).
Source: The authors (2024).
Source: The authors (2024).