Skip to content

Language Model¤

The language_model module is a wrapper around spaCy's training workflow for fine-tuning language models on custom corpora. It handles directory setup, config generation (including fine-tuning via component sourcing and transformer-based training via recipes), data conversion, training, evaluation, and packaging without requiring the user to edit config files or use the command line.

For a user-friendly overview, see Training Language Models in the User Guide. For hands-on walkthroughs, see the tutorial notebooks listed in Tutorials.

Constants¤

FULL_UD_PIPELINE: list[str] = ['tok2vec', 'tagger', 'morphologizer', 'trainable_lemmatizer', 'parser'] module-attribute ¤

rendering:
  show_root_heading: true
  heading_level: 3

CONLL-U Utilities¤

Standalone functions for preparing and managing CONLL-U training data.

split_conllu(input_path: str | Path, output_dir: str | Path, *, train_ratio: float = 0.8, dev_ratio: float = 0.1, seed: int = 42, shuffle: bool = True, include_test: bool = True) -> dict[str, Path] ¤

Split a single CONLL-U file into train / dev / (optionally) test files.

Sentences are the unit of splitting — sentence boundaries are blank lines in the CONLL-U format.

Parameters:

Name Type Description Default
input_path str | Path

Path to the source CONLL-U file.

required
output_dir str | Path

Directory where split files will be written.

required
train_ratio float

Fraction of sentences for the training split (default 0.8).

0.8
dev_ratio float

Fraction for the dev split (default 0.1). The test split receives whatever remains (1 - train_ratio - dev_ratio).

0.1
seed int

Random seed for reproducible shuffling (default 42).

42
shuffle bool

Whether to shuffle sentences before splitting. Set to False to preserve document order (e.g. split by act / chapter).

True
include_test bool

Whether to write a test file. Set to False for workflows that evaluate manually or with an external test set.

True

Returns:

Type Description
dict[str, Path]

Dict with keys "train", "dev", and (if include_test) "test"

dict[str, Path]

mapping to the Path of each written file. The return value is designed

dict[str, Path]

to be unpacked directly into :meth:LanguageModel.copy_assets::

splits = split_conllu("corpus.conllu", "model/assets/en/") model.copy_assets(**splits)

Source code in lexos/language_model/__init__.py
def split_conllu(
    input_path: str | Path,
    output_dir: str | Path,
    *,
    train_ratio: float = 0.8,
    dev_ratio: float = 0.1,
    seed: int = 42,
    shuffle: bool = True,
    include_test: bool = True,
) -> dict[str, Path]:
    """Split a single CONLL-U file into train / dev / (optionally) test files.

    Sentences are the unit of splitting — sentence boundaries are blank lines
    in the CONLL-U format.

    Args:
        input_path: Path to the source CONLL-U file.
        output_dir: Directory where split files will be written.
        train_ratio: Fraction of sentences for the training split (default 0.8).
        dev_ratio: Fraction for the dev split (default 0.1).  The test split
            receives whatever remains (1 - train_ratio - dev_ratio).
        seed: Random seed for reproducible shuffling (default 42).
        shuffle: Whether to shuffle sentences before splitting.  Set to False
            to preserve document order (e.g. split by act / chapter).
        include_test: Whether to write a test file.  Set to False for workflows
            that evaluate manually or with an external test set.

    Returns:
        Dict with keys `"train"`, `"dev"`, and (if include_test) `"test"`
        mapping to the Path of each written file.  The return value is designed
        to be unpacked directly into :meth:`LanguageModel.copy_assets`::

            splits = split_conllu("corpus.conllu", "model/assets/en/")
            model.copy_assets(**splits)
    """
    import random

    input_path = Path(input_path)
    output_dir = Path(output_dir)
    output_dir.mkdir(parents=True, exist_ok=True)

    text = input_path.read_text(encoding="utf-8")
    sentences = [s.strip() for s in text.split("\n\n") if s.strip()]

    if shuffle:
        rng = random.Random(seed)
        rng.shuffle(sentences)

    n = len(sentences)
    n_train = int(n * train_ratio)
    n_dev = int(n * dev_ratio)

    stem = input_path.stem
    splits: dict[str, list[str]] = {
        "train": sentences[:n_train],
        "dev": sentences[n_train : n_train + n_dev],
    }
    if include_test:
        splits["test"] = sentences[n_train + n_dev :]

    msg = Printer()
    paths: dict[str, Path] = {}
    for name, data in splits.items():
        out = output_dir / f"{stem}-{name}.conllu"
        out.write_text("\n\n".join(data) + "\n\n", encoding="utf-8")
        paths[name] = out
        msg.good(f"  {name:5s}: {len(data):4d} sentences  →  {out}")

    return paths
rendering:
  show_root_heading: true
  heading_level: 3

export_to_conllu(model_path: str | Path, texts: list[str], output_path: str | Path) -> Path ¤

Run a trained model on texts and write predictions to a CONLL-U file.

Each element of texts can be a single sentence or a longer passage; the model's sentence segmenter splits passages into individual sentences automatically.

Parameters:

Name Type Description Default
model_path str | Path

Path to a trained spaCy model directory, or an installed model name (e.g. "en_core_web_sm").

required
texts list[str]

List of strings to annotate. Each element may contain multiple sentences — the model's sentence segmenter handles splitting.

required
output_path str | Path

Path where the CONLL-U output file will be written.

required

Returns:

Type Description
Path

Path to the written output file.

Source code in lexos/language_model/__init__.py
def export_to_conllu(
    model_path: str | Path,
    texts: list[str],
    output_path: str | Path,
) -> Path:
    """Run a trained model on texts and write predictions to a CONLL-U file.

    Each element of `texts` can be a single sentence or a longer passage;
    the model's sentence segmenter splits passages into individual sentences
    automatically.

    Args:
        model_path: Path to a trained spaCy model directory, or an installed
            model name (e.g. `"en_core_web_sm"`).
        texts: List of strings to annotate.  Each element may contain multiple
            sentences — the model's sentence segmenter handles splitting.
        output_path: Path where the CONLL-U output file will be written.

    Returns:
        Path to the written output file.
    """
    # Using str() keeps installed-name inputs (e.g. "en_core_web_sm") loadable;
    # wrapping in Path() would force spaCy to treat them as directories.
    nlp = spacy.load(str(model_path))
    output_path = Path(output_path)
    msg = Printer()
    sent_id = 0
    with output_path.open("w", encoding="utf-8") as f:
        for text in texts:
            doc = nlp(text)
            for sent in doc.sents:
                sent_id = _write_sentence_conllu(f, sent, sent_id)
    msg.good(f"{sent_id} sentences written to {output_path}")
    return output_path
rendering:
  show_root_heading: true
  heading_level: 3

combine_conllu(round_files: list[str | Path], output_path: str | Path) -> Path ¤

Concatenate multiple CONLL-U files into a single training file.

Training on the full accumulated corpus each round (not just the latest batch) produces more stable models. Use this before each fine-tuning round to merge all corrected annotation batches.

Parameters:

Name Type Description Default
round_files list[str | Path]

List of paths to corrected CONLL-U files, in the order they should be concatenated.

required
output_path str | Path

Path where the combined output file will be written.

required

Returns:

Type Description
Path

Path to the written output file.

Source code in lexos/language_model/__init__.py
def combine_conllu(
    round_files: list[str | Path],
    output_path: str | Path,
) -> Path:
    """Concatenate multiple CONLL-U files into a single training file.

    Training on the full accumulated corpus each round (not just the latest
    batch) produces more stable models.  Use this before each fine-tuning
    round to merge all corrected annotation batches.

    Args:
        round_files: List of paths to corrected CONLL-U files, in the order
            they should be concatenated.
        output_path: Path where the combined output file will be written.

    Returns:
        Path to the written output file.
    """
    output_path = Path(output_path)
    msg = Printer()
    with output_path.open("w", encoding="utf-8") as out:
        for filepath in round_files:
            text = Path(filepath).read_text(encoding="utf-8")
            out.write(text)
            if not text.endswith("\n\n"):
                out.write("\n")
    msg.good(f"Combined {len(round_files)} file(s) → {output_path}")
    return output_path
rendering:
  show_root_heading: true
  heading_level: 3

The LanguageModel Class¤

The main entry point for the module is LanguageModel, a Pydantic-based configuration object that manages the model directory, training config, and workflow lifecycle. It is intended to be used as follows:

from lexos.language_model import LanguageModel

model = LanguageModel(
    model_dir="./model",
    lang="en",
    gpu=False,
    components=["tok2vec", "tagger", "morphologizer", "trainable_lemmatizer", "parser"],
)
model.copy_assets(train="train.conllu", dev="dev.conllu")
model.convert_assets()
model.train()

Public workflow methods exposed by the class include:

  • LanguageModel.copy_assets()
  • LanguageModel.convert_assets()
  • LanguageModel.validate()
  • LanguageModel.train()
  • LanguageModel.evaluate()
  • LanguageModel.package()
  • LanguageModel.config_path
  • LanguageModel.save_config()
  • LanguageModel.load_config()

This page intentionally avoids rendering the full Pydantic class with mkdocstrings because the generated schema for the model field defaults is not currently compatible with the installed griffe-pydantic template stack. The public methods above are the stable contract for the training workflow.

Debugging Utilities¤

Wrappers around spaCy's debugging commands for inspecting a model's config and data before training.

debug_config(config_path: str | Path, *, overrides: dict[str, Any] | None = None, code_path: str | Path | None = None, show_funcs: bool = False, show_vars: bool = False) -> None ¤

Validate a spaCy config file and report any errors.

Creates all registered objects described by the config and checks that every function reference is resolvable. Note: some validation errors are blocking — you may need to fix errors one round at a time.

Parameters:

Name Type Description Default
config_path str | Path

Path to the .cfg file to validate.

required
overrides dict[str, Any] | None

Dict of config key overrides to test (e.g. {"training.max_steps": 100}).

None
code_path str | Path | None

Path to a Python file containing custom registered functions.

None
show_funcs bool

Print all registered functions used by the config.

False
show_vars bool

Print all config variables and their resolved values.

False
Source code in lexos/language_model/__init__.py
def debug_config(
    config_path: str | Path,
    *,
    overrides: dict[str, Any] | None = None,
    code_path: str | Path | None = None,
    show_funcs: bool = False,
    show_vars: bool = False,
) -> None:
    """Validate a spaCy config file and report any errors.

    Creates all registered objects described by the config and checks that
    every function reference is resolvable.  Note: some validation errors are
    blocking — you may need to fix errors one round at a time.

    Args:
        config_path: Path to the `.cfg` file to validate.
        overrides: Dict of config key overrides to test (e.g.
            `{"training.max_steps": 100}`).
        code_path: Path to a Python file containing custom registered functions.
        show_funcs: Print all registered functions used by the config.
        show_vars: Print all config variables and their resolved values.
    """
    if overrides is None:
        overrides = {}
    config_path = Path(config_path)
    if isinstance(code_path, str):
        code_path = Path(code_path)
    import_code(code_path)
    try:
        spacy_debug_config(
            config_path, overrides=overrides, show_funcs=show_funcs, show_vars=show_vars
        )
    except SystemExit as e:
        if e.code != 0:
            raise LexosException(
                "debug_config found errors in your config (see output above). "
                "Fix them before training."
            ) from e
rendering:
  show_root_heading: true
  heading_level: 3

debug_data(config_path: str | Path, *, overrides: dict[str, Any] | None = None, code_path: str | Path | None = None, ignore_warnings: bool = False, verbose: bool = False, no_format: bool = False) -> None ¤

Analyse and validate training and dev data, reporting stats and issues.

Useful for catching problems like missing labels, data imbalance, or invalid annotations before a long training run. Raises LexosException if spaCy's data checker finds errors.

Parameters:

Name Type Description Default
config_path str | Path

Path to the .cfg file that references the data.

required
overrides dict[str, Any] | None

Dict of config key overrides.

None
code_path str | Path | None

Path to a Python file with custom registered functions.

None
ignore_warnings bool

Show only errors, not warnings.

False
verbose bool

Print additional explanations alongside stats.

False
no_format bool

Plain-text output without colour formatting.

False
Source code in lexos/language_model/__init__.py
def debug_data(
    config_path: str | Path,
    *,
    overrides: dict[str, Any] | None = None,
    code_path: str | Path | None = None,
    ignore_warnings: bool = False,
    verbose: bool = False,
    no_format: bool = False,
) -> None:
    """Analyse and validate training and dev data, reporting stats and issues.

    Useful for catching problems like missing labels, data imbalance, or
    invalid annotations before a long training run.  Raises `LexosException`
    if spaCy's data checker finds errors.

    Args:
        config_path: Path to the `.cfg` file that references the data.
        overrides: Dict of config key overrides.
        code_path: Path to a Python file with custom registered functions.
        ignore_warnings: Show only errors, not warnings.
        verbose: Print additional explanations alongside stats.
        no_format: Plain-text output without colour formatting.
    """
    if overrides is None:
        overrides = {}
    config_path = Path(config_path)
    if isinstance(code_path, str):
        code_path = Path(code_path)
    import_code(code_path)
    try:
        spacy_debug_data(
            config_path,
            config_overrides=overrides,
            ignore_warnings=ignore_warnings,
            verbose=verbose,
            no_format=no_format,
            silent=False,
        )
    except SystemExit as e:
        if e.code != 0:
            raise LexosException(
                "debug_data found errors in your training data (see output above). "
                "Fix them before training."
            ) from e
rendering:
  show_root_heading: true
  heading_level: 3

debug_model(config_path: str | Path, *, config_overrides: dict[str, Any] | None = None, component: str = 'tagger', layers: list[int] | None = None, dimensions: bool = False, parameters: bool = False, gradients: bool = False, attributes: bool = False, P0: bool = False, P1: bool = False, P2: bool = False, P3: bool = False, use_gpu: int = -1) -> None ¤

Inspect a trained model's internal layer structure and weights.

Parameters:

Name Type Description Default
config_path str | Path

Path to the .cfg file.

required
config_overrides dict[str, Any] | None

Dict of config key overrides.

None
component str

Pipeline component to inspect (default "tagger").

'tagger'
layers list[int] | None

Layer IDs to examine in detail.

None
dimensions bool

Print layer dimensions.

False
parameters bool

Print parameter counts.

False
gradients bool

Print gradient information.

False
attributes bool

Print component attributes.

False
P0 bool

Print model state before training.

False
P1 bool

Print model state after initialisation.

False
P2 bool

Print model state after training.

False
P3 bool

Print final predictions.

False
use_gpu int

GPU device ID or -1 for CPU.

-1
Source code in lexos/language_model/__init__.py
def debug_model(
    config_path: str | Path,
    *,
    config_overrides: dict[str, Any] | None = None,
    component: str = "tagger",
    layers: list[int] | None = None,
    dimensions: bool = False,
    parameters: bool = False,
    gradients: bool = False,
    attributes: bool = False,
    P0: bool = False,
    P1: bool = False,
    P2: bool = False,
    P3: bool = False,
    use_gpu: int = -1,
) -> None:
    """Inspect a trained model's internal layer structure and weights.

    Args:
        config_path: Path to the `.cfg` file.
        config_overrides: Dict of config key overrides.
        component: Pipeline component to inspect (default `"tagger"`).
        layers: Layer IDs to examine in detail.
        dimensions: Print layer dimensions.
        parameters: Print parameter counts.
        gradients: Print gradient information.
        attributes: Print component attributes.
        P0: Print model state before training.
        P1: Print model state after initialisation.
        P2: Print model state after training.
        P3: Print final predictions.
        use_gpu: GPU device ID or `-1` for CPU.
    """
    if config_overrides is None:
        config_overrides = {}
    if layers is None:
        layers = []
    config_path = Path(config_path)
    setup_gpu(use_gpu)
    print_settings = {
        "dimensions": dimensions,
        "parameters": parameters,
        "gradients": gradients,
        "attributes": attributes,
        "layers": [int(x) for x in layers],
        "print_before_training": P0,
        "print_after_init": P1,
        "print_after_training": P2,
        "print_prediction": P3,
    }
    with show_validation_error(config_path):
        raw_config = load_config(
            config_path, overrides=config_overrides, interpolate=False
        )
    config = raw_config.interpolate()
    allocator = config["training"]["gpu_allocator"]
    if use_gpu >= 0 and allocator:
        set_gpu_allocator(allocator)
    with show_validation_error(config_path):
        nlp = load_model_from_config(raw_config)
        config = nlp.config.interpolate()
        T = registry.resolve(config["training"], schema=ConfigSchemaTraining)
    seed = T["seed"]
    if seed is not None:
        fix_random_seed(seed)
    pipe = nlp.get_pipe(component)
    spacy_debug_model(config, T, nlp, pipe, print_settings=print_settings)
rendering:
  show_root_heading: true
  heading_level: 3

fill_config(config_path: str | Path, output_file: str | Path, *, pretraining: bool = False, diff: bool = False, code_path: str | Path | None = None) -> None ¤

Fill a partial config file with spaCy defaults and save it.

Useful for debugging or understanding what a minimal config expands to.

Parameters:

Name Type Description Default
config_path str | Path

Path to the partial .cfg file to fill.

required
output_file str | Path

Path where the filled config will be written.

required
pretraining bool

Include pretraining config section.

False
diff bool

Print a visual diff of changes made.

False
code_path str | Path | None

Path to a Python file with custom registered functions.

None
Source code in lexos/language_model/__init__.py
def fill_config(
    config_path: str | Path,
    output_file: str | Path,
    *,
    pretraining: bool = False,
    diff: bool = False,
    code_path: str | Path | None = None,
) -> None:
    """Fill a partial config file with spaCy defaults and save it.

    Useful for debugging or understanding what a minimal config expands to.

    Args:
        config_path: Path to the partial `.cfg` file to fill.
        output_file: Path where the filled config will be written.
        pretraining: Include pretraining config section.
        diff: Print a visual diff of changes made.
        code_path: Path to a Python file with custom registered functions.
    """
    config_path = Path(config_path)
    output_file = Path(output_file)
    if isinstance(code_path, str):
        code_path = Path(code_path)
    import_code(code_path)
    spacy_fill_config(output_file, config_path, pretraining=pretraining, diff=diff)
rendering:
  show_root_heading: true
  heading_level: 3

Internal Helpers¤

The following helpers are used internally by the training workflow and are not part of the public API contract:

  • _has_nvidia_gpu()
  • _get_tok2vec_width()
  • _patch_tok2vec_width()