API Reference#

The FastAPI application called “app”.

Application#

Routes#

Core#

Assessment data model#

Bookkeeping for research-file imports.

The importer is expected to run repeatedly against the same object store prefix, so it records what it has already loaded. A source object is skipped when its size and entity tag still match the row written by the last successful load, which is what lets a scheduled run in AWS re-read the whole bucket and do work only for the files that actually changed.

class app.model.ingest.IngestFile(*, id: ~typing.Annotated[~uuid.UUID, ~sqlmodel.main.FieldInfoMetadata(primary_key=True, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = <factory>, runId: ~typing.Annotated[~uuid.UUID, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=ingest_runs.id, ondelete=PydanticUndefined, unique=PydanticUndefined, index=True, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)], sourceKey: ~typing.Annotated[str, ~annotated_types.MaxLen(max_length=500), ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)], etag: ~typing.Annotated[str | None, ~annotated_types.MaxLen(max_length=200), ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, sizeBytes: ~typing.Annotated[int | None, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, lastModified: ~typing.Annotated[~datetime.datetime | None, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, program: ~typing.Annotated[str | None, ~annotated_types.MaxLen(max_length=10), ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, testType: ~typing.Annotated[str | None, ~annotated_types.MaxLen(max_length=10), ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, testYear: ~typing.Annotated[int | None, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, status: ~typing.Annotated[~app.model.ingest.IngestStatus, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=Enum('running', 'succeeded', 'failed', 'skipped', name='ingeststatus'), sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = IngestStatus.RUNNING, resultRows: ~typing.Annotated[int, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = 0, subscoreRows: ~typing.Annotated[int, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = 0, durationSeconds: ~typing.Annotated[float | None, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, error: ~typing.Annotated[str | None, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, loadedAt: ~typing.Annotated[~datetime.datetime | None, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None)#

Bases: ApiModel

The outcome of loading a single research file.

duration_seconds: float | None#
error: str | None#
etag: str | None#
id: uuid.UUID#
last_modified: datetime | None#
loaded_at: datetime | None#
model_config = {'alias_generator': <function to_camel>, 'from_attributes': True, 'populate_by_name': True, 'read_from_attributes': True, 'read_with_orm_mode': True, 'registry': PydanticUndefined, 'table': True, 'validate_by_alias': True, 'validate_by_name': True}#

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

program: str | None#
result_rows: int#
run_id: uuid.UUID#
size_bytes: int | None#
source_key: str#
status: IngestStatus#
subscore_rows: int#
test_type: str | None#
test_year: int | None#
class app.model.ingest.IngestRun(*, id: ~typing.Annotated[~uuid.UUID, ~sqlmodel.main.FieldInfoMetadata(primary_key=True, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = <factory>, sourceUri: ~typing.Annotated[str, ~annotated_types.MaxLen(max_length=500), ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)], status: ~typing.Annotated[~app.model.ingest.IngestStatus, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=Enum('running', 'succeeded', 'failed', 'skipped', name='ingeststatus'), sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = IngestStatus.RUNNING, startedAt: ~datetime.datetime, finishedAt: ~typing.Annotated[~datetime.datetime | None, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None, filesSeen: ~typing.Annotated[int, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = 0, filesLoaded: ~typing.Annotated[int, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = 0, filesSkipped: ~typing.Annotated[int, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = 0, resultRows: ~typing.Annotated[int, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = 0, subscoreRows: ~typing.Annotated[int, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = 0, error: ~typing.Annotated[str | None, ~sqlmodel.main.FieldInfoMetadata(primary_key=PydanticUndefined, nullable=PydanticUndefined, foreign_key=PydanticUndefined, ondelete=PydanticUndefined, unique=PydanticUndefined, index=PydanticUndefined, sa_type=PydanticUndefined, sa_column=PydanticUndefined, sa_column_args=PydanticUndefined, sa_column_kwargs=PydanticUndefined)] = None)#

Bases: ApiModel

One invocation of the importer over a source location.

error: str | None#
files_loaded: int#
files_seen: int#
files_skipped: int#
finished_at: datetime | None#
id: uuid.UUID#
model_config = {'alias_generator': <function to_camel>, 'from_attributes': True, 'populate_by_name': True, 'read_from_attributes': True, 'read_with_orm_mode': True, 'registry': PydanticUndefined, 'table': True, 'validate_by_alias': True, 'validate_by_name': True}#

Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].

result_rows: int#
source_uri: str#
started_at: datetime#
status: IngestStatus#
subscore_rows: int#
class app.model.ingest.IngestStatus(*values)#

Bases: StrEnum

Lifecycle of an ingest run or of one file within it.

FAILED = 'failed'#
RUNNING = 'running'#
SKIPPED = 'skipped'#
SUCCEEDED = 'succeeded'#

Research file importer#

Import CAASPP and ELPAC research files into the application database.

The pipeline has four pieces:

app.ingest.sources

Where files come from – a local directory or an S3 prefix.

app.ingest.layouts

Which columns each published research file has, keyed by test type and administration year.

app.ingest.parser

Row-level conversion, including the state’s two kinds of missing value.

app.ingest.loader

Bulk COPY into staging tables and an atomic swap into place.

app.ingest.runner.ImportRunner ties them together and records what was loaded so repeated runs only do new work.

class app.ingest.FileOutcome(name, status, layout_key=None, test_year=None, results=0, subscores=0, duration_seconds=0.0, error=None)#

Bases: object

What happened to one object during a run.

duration_seconds: float#
error: str | None#
layout_key: str | None#
name: str#
results: int#
status: IngestStatus#
subscores: int#
test_year: int | None#
class app.ingest.ImportRunner(engine)#

Bases: object

Runs imports and records what it did.

run(source_uri, *, force=False, only=None, years=None)#

Import every changed research file under source_uri.

Return type:

RunOutcome

Args:

source_uri: A local directory or an s3://bucket/prefix URI. force: Reload files even when their fingerprint is unchanged. only: Substrings; when given, only matching file names are loaded. years: Administration years to restrict the run to.

class app.ingest.LocalSource(root, *, encoding='cp1252')#

Bases: object

Reads research files from a directory tree.

list_objects()#
Return type:

Iterator[SourceObject]

open_file(obj)#
Return type:

Iterator[Path]

open_text(obj)#
Return type:

Iterator[Iterator[str]]

class app.ingest.RunOutcome(run_id, source_uri, status, files=<factory>)#

Bases: object

The result of a whole run.

files: list[FileOutcome]#
property results: int#
run_id: str#
source_uri: str#
status: IngestStatus#
property subscores: int#
class app.ingest.S3Source(bucket, prefix='', *, client=None, encoding='cp1252')#

Bases: object

Reads research files from an S3 bucket prefix.

property client: Any#
list_objects()#
Return type:

Iterator[SourceObject]

open_file(obj)#
Return type:

Iterator[Path]

open_text(obj)#
Return type:

Iterator[Iterator[str]]

app.ingest.source_from_uri(uri, *, encoding=None)#

Build a source from a path, an s3:// prefix or an https:// URL.

encoding overrides each source’s default. It is what lets the Dashboard families read a copy of their files out of a bucket or a directory: those are published as UTF-8, while the research files the other sources default to are code page 1252.

Return type:

ResearchFileSource

Where research files are read from.

Two sources are supported and they present the same interface, so the importer does not care which one it is given:

LocalSource

A directory tree, used for development and for one-off loads from a downloaded copy of the files.

S3Source

An S3 bucket and prefix, which is how the deployed application gets its data. New administrations are published by uploading them to the bucket; nothing has to be parsed or pushed from a workstation.

Both yield objects with the size, entity tag and modification time the importer records so it can skip files it has already loaded.

Research files are distributed as ZIP archives and are sometimes stored compressed, so .zip and .gz are unwrapped transparently. ZIP archives need random access, which an S3 response body cannot provide, so those are staged to a temporary file first.

class app.ingest.sources.HttpSource(base_url='https://www3.cde.ca.gov/researchfiles/cadashboard/', *, years=None, stems=('ela', 'math', 'chronic', 'susp', 'grad', 'elpi', 'cci', 'science'), extra_stems=('dass1yeargraduationrate', 'elpacpart'), encoding='utf-8', names=None, opener=None)#

Bases: object

Reads Dashboard indicator files straight from the state’s web server.

Unlike a bucket or a directory there is nothing to list, so the candidate file names are generated from the published naming convention – {indicator}download{year}.txt – and a HEAD request decides whether each one exists and whether it has changed. The state revises these files in place after release, which is exactly what the entity tag catches.

list_objects()#
Return type:

Iterator[SourceObject]

open_file(obj)#
Return type:

Iterator[Path]

open_text(obj)#
Return type:

Iterator[Iterator[str]]

class app.ingest.sources.LocalSource(root, *, encoding='cp1252')#

Bases: object

Reads research files from a directory tree.

list_objects()#
Return type:

Iterator[SourceObject]

open_file(obj)#
Return type:

Iterator[Path]

open_text(obj)#
Return type:

Iterator[Iterator[str]]

class app.ingest.sources.ResearchFileSource(*args, **kwargs)#

Bases: Protocol

A place research files can be read from.

encoding: str#
list_objects()#

Yield every candidate research file, in a stable order.

Return type:

Iterator[SourceObject]

open_file(obj)#

Make an object available as a real file on disk.

Needed by readers that seek – a spreadsheet, for instance – which a streamed HTTP or S3 response body cannot support.

Return type:

AbstractContextManager[Path]

open_text(obj)#

Open an object as decoded text lines.

Return type:

AbstractContextManager[Iterator[str]]

uri: str#
class app.ingest.sources.S3Source(bucket, prefix='', *, client=None, encoding='cp1252')#

Bases: object

Reads research files from an S3 bucket prefix.

property client: Any#
list_objects()#
Return type:

Iterator[SourceObject]

open_file(obj)#
Return type:

Iterator[Path]

open_text(obj)#
Return type:

Iterator[Iterator[str]]

class app.ingest.sources.SourceObject(key, name, size_bytes=None, etag=None, last_modified=None)#

Bases: object

One candidate research file in a source location.

etag: str | None#
property fingerprint: str#

A value that changes whenever the object’s contents change.

key: str#
last_modified: datetime | None#
name: str#
size_bytes: int | None#
app.ingest.sources.is_delimited_text(name)#

Whether a file is one the delimited text readers can parse.

The Dashboard families are published as tab separated text, but the state also publishes spreadsheets – the Local Indicators are only available that way – and a bucket mirroring both has them side by side. Handing a spreadsheet to a text reader decodes the ZIP container as text, so the importers use this to leave files that are not theirs alone.

An archive counts, since it unwraps to a single data member.

Return type:

bool

app.ingest.sources.source_from_uri(uri, *, encoding=None)#

Build a source from a path, an s3:// prefix or an https:// URL.

encoding overrides each source’s default. It is what lets the Dashboard families read a copy of their files out of a bucket or a directory: those are published as UTF-8, while the research files the other sources default to are code page 1252.

Return type:

ResearchFileSource

Declarative descriptions of the state’s research file layouts.

Every research file is a delimited table with a header row, so a layout is expressed as the names of the columns to read rather than their positions. That keeps the importer working when the state adds a column mid-file – which it does – and makes a mismatch fail loudly at load time instead of silently shifting every value by one.

Two conventions have to be reconciled here:

  • CAASPP files use spaced title case (Student Group ID) and carry a Test Type column; ELPAC files use compact camel case (StudentGroupID) and identify the test through TestID alone.

  • Performance bands are printed in different directions. Smarter Balanced areas run above/near/below standard while CAST domains run below/near/above. Each layout lists its bands lowest-first, and the importer stores them in that order, so downstream queries never have to know which file a row came from.

Sources: the fixed-length record definition pages at https://caaspp-elpac.ets.org/caaspp/ResearchFileFormat{SB,CAA,CAST,CAAS,CSA} and https://caaspp-elpac.ets.org/elpac/ResearchFileFormat{SA,IA,ALTSA,ALTIA}, each of which also publishes the caret-delimited column header for every field.

class app.ingest.layouts.BandColumns(pct, count)#

Bases: object

The percentage and count columns for one performance band.

count: str | None#
pct: str | None#
exception app.ingest.layouts.LayoutError#

Bases: RuntimeError

Raised when a file cannot be matched to a known research file layout.

class app.ingest.layouts.ResearchFileLayout(key, program, test_type, test_ids, min_year=2015, max_year=9999, county_code='County Code', district_code='District Code', school_code='School Code', type_id='Type ID', test_year='Test Year', test_id='Test ID', student_group_id='Student Group ID', grade='Grade', county_name=None, district_name='District Name', school_name='School Name', zip_code=None, students_enrolled='Total Students Enrolled', students_tested='Total Students Tested', students_tested_with_scores='Total Students Tested with Scores', mean_scale_score='Mean Scale Score', levels=(), met_or_above=None, overall_total='Overall Total', subscores=<factory>)#

Bases: object

How to read one family of research files.

county_code: str#
county_name: str | None#
covers_year(year)#
Return type:

bool

district_code: str#
district_name: str | None#
grade: str#
key: str#
levels: tuple[BandColumns, ...]#
max_year: int#
mean_scale_score: str | None#
met_or_above: BandColumns | None#
min_year: int#
overall_total: str | None#
program: Program#
property required_columns: tuple[str, ...]#

Columns that must be present for the layout to match a header.

school_code: str#
school_name: str | None#
student_group_id: str#
students_enrolled: str#
students_tested: str#
students_tested_with_scores: str#
subscores: tuple[SubscoreColumns, ...]#
test_id: str#
test_ids: tuple[int, ...]#
test_type: str#
test_year: str#
type_id: str#
zip_code: str | None#
class app.ingest.layouts.SubscoreColumns(code, bands, total=None, mean_scale_score=None)#

Bases: object

Where one area, domain or composite lives in a research file.

bands: tuple[BandColumns, ...]#
code: str#
mean_scale_score: str | None#
total: str | None#
app.ingest.layouts.candidate_keys(name)#

Layout keys suggested by a file name, most specific first.

Return type:

tuple[str, ...]

app.ingest.layouts.resolve_layout(name, header, *, test_year=None)#

Choose the layout that matches a file’s name, header and year.

The name only narrows the search; the header decides. A file whose name is unrecognised still loads as long as its columns match exactly one layout, which is what makes renamed or hand-extracted files work.

Return type:

ResearchFileLayout

app.ingest.layouts.year_from_filename(name)#

Pull the administration year out of a research file name.

Return type:

int | None

Turn research file rows into database records.

The state encodes two different kinds of “no value” and they mean different things, so both survive the trip into the database:

An asterisk means the figure exists but is withheld, because the group is too small to report without identifying students; the row is flagged suppressed. An empty field means the figure does not apply – most often the mean scale score on an “all grades” row, which the state deliberately leaves blank because scale scores are not comparable between grades.

Both land as NULL; only the first sets the flag.

class app.ingest.parser.EntityRecord(cds_code, county_code, district_code, school_code, entity_level, type_id, is_charter, charter_funding, county_name, district_name, school_name, zip_code, display_name, parent_cds_code, first_test_year, last_test_year)#

Bases: object

An entity discovered while reading a research file.

cds_code: str#
charter_funding: CharterFunding | None#
county_code: str#
county_name: str | None#
display_name: str#
district_code: str#
district_name: str | None#
entity_level: EntityLevel#
first_test_year: int | None#
is_charter: bool#
last_test_year: int | None#
parent_cds_code: str | None#
school_code: str#
school_name: str | None#
type_id: int#
zip_code: str | None#
exception app.ingest.parser.ParseError#

Bases: ValueError

Raised when a row cannot be interpreted with the resolved layout.

class app.ingest.parser.ParsedRows(result, subscores=<factory>, entity=None)#

Bases: object

Everything one research file row contributes to the database.

entity: EntityRecord | None#
result: ResultRecord#
subscores: list[SubscoreRecord]#
class app.ingest.parser.ResultRecord(cds_code, test_year, test_id, student_group_id, grade, students_enrolled, students_tested, students_tested_with_scores, mean_scale_score, level_counts, level_pcts, met_or_above_count, met_or_above_pct, met_or_above_source, overall_total, suppressed)#

Bases: object

One overall result cell, ready for AssessmentResult.

cds_code: str#
grade: str#
level_counts: list[int | None]#
level_pcts: list[Decimal | None]#
mean_scale_score: Decimal | None#
met_or_above_count: int | None#
met_or_above_pct: Decimal | None#
met_or_above_source: MetOrAboveSource | None#
overall_total: int | None#
student_group_id: int#
students_enrolled: int | None#
students_tested: int | None#
students_tested_with_scores: int | None#
suppressed: bool#
test_id: int#
test_year: int#
class app.ingest.parser.RowParser(layout, header)#

Bases: object

Reads rows of one research file using a resolved layout.

parse(row, *, default_year=None)#

Convert one row into its result, subscore and entity records.

Return type:

ParsedRows

class app.ingest.parser.SubscoreRecord(cds_code, test_year, test_id, student_group_id, grade, subscore_code, mean_scale_score, band_counts, band_pcts, subscore_total)#

Bases: object

One area, domain or composite breakdown.

band_counts: list[int | None]#
band_pcts: list[Decimal | None]#
cds_code: str#
grade: str#
mean_scale_score: Decimal | None#
student_group_id: int#
subscore_code: str#
subscore_total: int | None#
test_id: int#
test_year: int#
app.ingest.parser.build_cds_code(county, district, school)#

Assemble the 14-character CDS code from its three parts.

Return type:

str

app.ingest.parser.entity_level_for(county, district, school)#

Derive the reporting level from the code parts.

The programs number their Type ID values differently – CAASPP uses 07 for a school where ELPAC uses 01 – so the codes, which agree, are the reliable signal.

Return type:

EntityLevel

app.ingest.parser.iter_rows(lines, delimiter)#

Split already-decoded lines on the file’s delimiter.

Research files use a caret or, for administrations through 2018-19, a comma. Neither is ever quoted, so a plain split is both correct and considerably faster than the csv module.

Return type:

Iterator[list[str]]

app.ingest.parser.normalize_grade(raw)#

Normalise a grade code to the two characters used in the layouts.

Return type:

str

app.ingest.parser.parent_cds_for(level, county, district)#

The CDS code of the entity one level up, or None for the state.

Return type:

str | None

app.ingest.parser.parse_decimal(raw)#

Read a decimal, treating suppressed and empty values as missing.

Return type:

Decimal | None

app.ingest.parser.parse_int(raw)#

Read an integer, treating suppressed and empty values as missing.

Return type:

int | None

Bulk-load parsed research file rows into PostgreSQL.

A statewide research file is a few million rows, so rows go in through COPY into an unlogged staging table and are then swapped into place in one transaction. Loading a file is idempotent: everything already stored for the test years and test IDs the file covers is deleted and replaced, so re-running the importer over an unchanged bucket converges rather than duplicating.

Entities are collected as a side effect of reading the rows. Names arrive piecemeal – a school row carries only the school and district names, a district row only the district name – so entity columns are merged rather than overwritten, and the years an entity appears in are widened, never narrowed.

class app.ingest.loader.LoadCounts(results=0, subscores=0, entities=0, test_years=None, test_ids=None)#

Bases: object

How much a single file contributed.

entities: int#
results: int#
subscores: int#
test_ids: set[int] | None#
test_years: set[int] | None#
class app.ingest.loader.ResearchFileLoader(engine)#

Bases: object

Loads one research file’s parsed rows into the database.

load(rows)#

Stage, then atomically replace, everything a file covers.

Return type:

LoadCounts

app.ingest.loader.analyze(engine)#

Refresh planner statistics after a load.

Return type:

None

Orchestrates an import of research files from a source location.

A run walks every object under the source, matches each one to a published research file layout, and loads the ones that have changed since the last successful run. What “changed” means is the object’s size and entity tag, so pointing the importer at an S3 prefix and running it on a schedule loads only newly uploaded administrations.

Reference data is seeded first, because result rows reference the assessment catalogue and the parser needs the county names.

class app.ingest.runner.FileOutcome(name, status, layout_key=None, test_year=None, results=0, subscores=0, duration_seconds=0.0, error=None)#

Bases: object

What happened to one object during a run.

duration_seconds: float#
error: str | None#
layout_key: str | None#
name: str#
results: int#
status: IngestStatus#
subscores: int#
test_year: int | None#
class app.ingest.runner.ImportRunner(engine)#

Bases: object

Runs imports and records what it did.

run(source_uri, *, force=False, only=None, years=None)#

Import every changed research file under source_uri.

Return type:

RunOutcome

Args:

source_uri: A local directory or an s3://bucket/prefix URI. force: Reload files even when their fingerprint is unchanged. only: Substrings; when given, only matching file names are loaded. years: Administration years to restrict the run to.

class app.ingest.runner.RunOutcome(run_id, source_uri, status, files=<factory>)#

Bases: object

The result of a whole run.

files: list[FileOutcome]#
property results: int#
run_id: str#
source_uri: str#
status: IngestStatus#
property subscores: int#
app.ingest.runner.detect_delimiter(header_line)#

Pick the delimiter a research file uses from its header row.

Return type:

str

Seed data for the CAASPP and ELPAC reference tables.

Everything here is transcribed from the state’s published record definitions and “Understanding Results” pages, which are the authoritative source for the vocabulary the research files only encode positionally:

  • Record layouts and Tables A/B/C – https://caaspp-elpac.ets.org/caaspp/ResearchFileFormatSB (and the CAA/CAST/CAAS/CSA variants), plus https://caaspp-elpac.ets.org/elpac/ResearchFileFormatSA (and IA, ALTSA, ALTIA).

  • Achievement level descriptors – https://caaspp-elpac.ets.org/caaspp/UnderstandingCAAResults, .../UnderstandingCSAResults, https://caaspp-elpac.ets.org/elpac/UnderstandingReportsSA and .../UnderstandingReportsIA.

Seeding is idempotent: rows are upserted on their primary keys so a redeploy can safely re-run it.

app.ingest.reference_data.level_scheme_for(test_id, test_year)#

Public alias for the year-aware achievement level scheme lookup.

Return type:

str

app.ingest.reference_data.proficient_from_level(test_id, test_year)#

The lowest level a test counts as meeting the standard, if any.

Used to derive a “met or above” figure for the tests whose research files do not publish one; Smarter Balanced and CAST publish it directly.

Return type:

int | None

app.ingest.reference_data.seed_reference_data(session)#

Populate every reference table. Safe to call repeatedly.

Return type:

None

app.ingest.reference_data.subscores_for(test_id, test_year)#

Return the subscores a test published in a given year.

Return type:

tuple[tuple[str, SubscoreKind, str, str, bool, int], ...]

Reporting services#

Data models and CRUD#

Utilities#