Language Model¤
The language_model module is a wrapper around spaCy's training workflow for fine-tuning language models on custom corpora. It handles directory setup, config generation (including fine-tuning via component sourcing and transformer-based training via recipes), data conversion, training, evaluation, and packaging without requiring the user to edit config files or use the command line.
For a user-friendly overview, see Training Language Models in the User Guide. For hands-on walkthroughs, see the tutorial notebooks listed in Tutorials.
Constants¤
FULL_UD_PIPELINE: list[str] = ['tok2vec', 'tagger', 'morphologizer', 'trainable_lemmatizer', 'parser']
module-attribute
¤
rendering:
show_root_heading: true
heading_level: 3
CONLL-U Utilities¤
Standalone functions for preparing and managing CONLL-U training data.
split_conllu(input_path: str | Path, output_dir: str | Path, *, train_ratio: float = 0.8, dev_ratio: float = 0.1, seed: int = 42, shuffle: bool = True, include_test: bool = True) -> dict[str, Path]
¤
Split a single CONLL-U file into train / dev / (optionally) test files.
Sentences are the unit of splitting — sentence boundaries are blank lines in the CONLL-U format.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_path
|
str | Path
|
Path to the source CONLL-U file. |
required |
output_dir
|
str | Path
|
Directory where split files will be written. |
required |
train_ratio
|
float
|
Fraction of sentences for the training split (default 0.8). |
0.8
|
dev_ratio
|
float
|
Fraction for the dev split (default 0.1). The test split receives whatever remains (1 - train_ratio - dev_ratio). |
0.1
|
seed
|
int
|
Random seed for reproducible shuffling (default 42). |
42
|
shuffle
|
bool
|
Whether to shuffle sentences before splitting. Set to False to preserve document order (e.g. split by act / chapter). |
True
|
include_test
|
bool
|
Whether to write a test file. Set to False for workflows that evaluate manually or with an external test set. |
True
|
Returns:
| Type | Description |
|---|---|
dict[str, Path]
|
Dict with keys |
dict[str, Path]
|
mapping to the Path of each written file. The return value is designed |
dict[str, Path]
|
to be unpacked directly into :meth: splits = split_conllu("corpus.conllu", "model/assets/en/") model.copy_assets(**splits) |
Source code in lexos/language_model/__init__.py
rendering:
show_root_heading: true
heading_level: 3
export_to_conllu(model_path: str | Path, texts: list[str], output_path: str | Path) -> Path
¤
Run a trained model on texts and write predictions to a CONLL-U file.
Each element of texts can be a single sentence or a longer passage;
the model's sentence segmenter splits passages into individual sentences
automatically.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_path
|
str | Path
|
Path to a trained spaCy model directory, or an installed
model name (e.g. |
required |
texts
|
list[str]
|
List of strings to annotate. Each element may contain multiple sentences — the model's sentence segmenter handles splitting. |
required |
output_path
|
str | Path
|
Path where the CONLL-U output file will be written. |
required |
Returns:
| Type | Description |
|---|---|
Path
|
Path to the written output file. |
Source code in lexos/language_model/__init__.py
rendering:
show_root_heading: true
heading_level: 3
combine_conllu(round_files: list[str | Path], output_path: str | Path) -> Path
¤
Concatenate multiple CONLL-U files into a single training file.
Training on the full accumulated corpus each round (not just the latest batch) produces more stable models. Use this before each fine-tuning round to merge all corrected annotation batches.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
round_files
|
list[str | Path]
|
List of paths to corrected CONLL-U files, in the order they should be concatenated. |
required |
output_path
|
str | Path
|
Path where the combined output file will be written. |
required |
Returns:
| Type | Description |
|---|---|
Path
|
Path to the written output file. |
Source code in lexos/language_model/__init__.py
rendering:
show_root_heading: true
heading_level: 3
The LanguageModel Class¤
The main entry point for the module is LanguageModel, a Pydantic-based configuration object that manages the model directory, training config, and workflow lifecycle. It is intended to be used as follows:
from lexos.language_model import LanguageModel
model = LanguageModel(
model_dir="./model",
lang="en",
gpu=False,
components=["tok2vec", "tagger", "morphologizer", "trainable_lemmatizer", "parser"],
)
model.copy_assets(train="train.conllu", dev="dev.conllu")
model.convert_assets()
model.train()
Public workflow methods exposed by the class include:
LanguageModel.copy_assets()LanguageModel.convert_assets()LanguageModel.validate()LanguageModel.train()LanguageModel.evaluate()LanguageModel.package()LanguageModel.config_pathLanguageModel.save_config()LanguageModel.load_config()
This page intentionally avoids rendering the full Pydantic class with mkdocstrings because the generated schema for the model field defaults is not currently compatible with the installed griffe-pydantic template stack. The public methods above are the stable contract for the training workflow.
Debugging Utilities¤
Wrappers around spaCy's debugging commands for inspecting a model's config and data before training.
debug_config(config_path: str | Path, *, overrides: dict[str, Any] | None = None, code_path: str | Path | None = None, show_funcs: bool = False, show_vars: bool = False) -> None
¤
Validate a spaCy config file and report any errors.
Creates all registered objects described by the config and checks that every function reference is resolvable. Note: some validation errors are blocking — you may need to fix errors one round at a time.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config_path
|
str | Path
|
Path to the |
required |
overrides
|
dict[str, Any] | None
|
Dict of config key overrides to test (e.g.
|
None
|
code_path
|
str | Path | None
|
Path to a Python file containing custom registered functions. |
None
|
show_funcs
|
bool
|
Print all registered functions used by the config. |
False
|
show_vars
|
bool
|
Print all config variables and their resolved values. |
False
|
Source code in lexos/language_model/__init__.py
rendering:
show_root_heading: true
heading_level: 3
debug_data(config_path: str | Path, *, overrides: dict[str, Any] | None = None, code_path: str | Path | None = None, ignore_warnings: bool = False, verbose: bool = False, no_format: bool = False) -> None
¤
Analyse and validate training and dev data, reporting stats and issues.
Useful for catching problems like missing labels, data imbalance, or
invalid annotations before a long training run. Raises LexosException
if spaCy's data checker finds errors.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config_path
|
str | Path
|
Path to the |
required |
overrides
|
dict[str, Any] | None
|
Dict of config key overrides. |
None
|
code_path
|
str | Path | None
|
Path to a Python file with custom registered functions. |
None
|
ignore_warnings
|
bool
|
Show only errors, not warnings. |
False
|
verbose
|
bool
|
Print additional explanations alongside stats. |
False
|
no_format
|
bool
|
Plain-text output without colour formatting. |
False
|
Source code in lexos/language_model/__init__.py
rendering:
show_root_heading: true
heading_level: 3
debug_model(config_path: str | Path, *, config_overrides: dict[str, Any] | None = None, component: str = 'tagger', layers: list[int] | None = None, dimensions: bool = False, parameters: bool = False, gradients: bool = False, attributes: bool = False, P0: bool = False, P1: bool = False, P2: bool = False, P3: bool = False, use_gpu: int = -1) -> None
¤
Inspect a trained model's internal layer structure and weights.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config_path
|
str | Path
|
Path to the |
required |
config_overrides
|
dict[str, Any] | None
|
Dict of config key overrides. |
None
|
component
|
str
|
Pipeline component to inspect (default |
'tagger'
|
layers
|
list[int] | None
|
Layer IDs to examine in detail. |
None
|
dimensions
|
bool
|
Print layer dimensions. |
False
|
parameters
|
bool
|
Print parameter counts. |
False
|
gradients
|
bool
|
Print gradient information. |
False
|
attributes
|
bool
|
Print component attributes. |
False
|
P0
|
bool
|
Print model state before training. |
False
|
P1
|
bool
|
Print model state after initialisation. |
False
|
P2
|
bool
|
Print model state after training. |
False
|
P3
|
bool
|
Print final predictions. |
False
|
use_gpu
|
int
|
GPU device ID or |
-1
|
Source code in lexos/language_model/__init__.py
1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 | |
rendering:
show_root_heading: true
heading_level: 3
fill_config(config_path: str | Path, output_file: str | Path, *, pretraining: bool = False, diff: bool = False, code_path: str | Path | None = None) -> None
¤
Fill a partial config file with spaCy defaults and save it.
Useful for debugging or understanding what a minimal config expands to.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
config_path
|
str | Path
|
Path to the partial |
required |
output_file
|
str | Path
|
Path where the filled config will be written. |
required |
pretraining
|
bool
|
Include pretraining config section. |
False
|
diff
|
bool
|
Print a visual diff of changes made. |
False
|
code_path
|
str | Path | None
|
Path to a Python file with custom registered functions. |
None
|
Source code in lexos/language_model/__init__.py
rendering:
show_root_heading: true
heading_level: 3
Internal Helpers¤
The following helpers are used internally by the training workflow and are not part of the public API contract:
_has_nvidia_gpu()_get_tok2vec_width()_patch_tok2vec_width()