Writing Conventions for Linguistics
Data appear as numbered examples referenced by number
Example sentences are set off, numbered sequentially as (1), (2), with sub-items lettered (3a), (3b), and the prose refers back to them rather than repeating the string. Every numbered example must be discussed somewhere in the text, or a reviewer will ask why it is there.
Non-English data are glossed line by line
The standard is a three-line format — the object language, a morpheme-by-morpheme gloss with grammatical categories in small caps, and a free translation in single quotes. The Leipzig Glossing Rules define the abbreviations and the alignment, and misaligned morpheme boundaries make the example unreadable.
Grammaticality is marked with a fixed set of diacritics
An asterisk marks ungrammaticality, a question mark marginal acceptability, a hash a semantic or pragmatic anomaly. These are conventional and load-bearing, so you say in the text where the judgments come from — your own intuitions, consultants, or an acceptability study.
Transcription notation distinguishes phonemic from phonetic
Slashes enclose phonemic representations and square brackets phonetic ones, with IPA symbols rather than ad-hoc respellings. Italics mark forms cited as forms, and single quotes enclose meanings — a distinction the field enforces closely.
Language data carry provenance and speaker ethics
For fieldwork data, state the language with its identifying code, where and when it was collected, and how consultants consented and are credited. Anonymization practices and community agreements are described in the methods, not assumed.
Typological claims are hedged to the sample they rest on
A generalization drawn from a convenience sample of a dozen languages is written as such. The field is attentive to the difference between “no language does X” and “X is unattested in our sample,” and the second is usually the honest sentence.