Blog

Co-creating Code with LLMs

| Colin Kerkhof

In AI engineering we keep looking for ways to use Large Language Models (LLMs) for more than code completion, and to bring them into the more complex, project-specific parts of development. Generating snippets is the easy part. The challenge is fitting the LLM into our development lifecycle as a proper collaborator. How do you guide one to build robust, project-specific boilerplate that follows your own architectural patterns?

On a recent project we built a system to help streamline the review of social security benefit applications. That work involves complex business rules, detailed evidence review and strict data consistency requirements. Writing the boilerplate for data models, APIs and tests by hand, for every type of benefit and condition, would have been slow and error-prone. It was a good candidate for our LLM co-creation workflow, which let us generate consistent code from a set of well-defined patterns.

This post describes a structured, iterative workflow that puts the engineer in the role of architect and the LLM in the role of a highly skilled pair programmer. The collaboration rests on a simple mantra: "Apply Pattern X to our specific Context Y." With clear patterns and focused context, the work shifts from writing boilerplate to designing systems.

This is a happy path that has worked well for us. Try it, and adapt it to your own way of working.

Our development workflow: a five-step cycle

We have distilled the process into five steps that take an idea from rough concept to an implemented and tested feature, all in close collaboration with an LLM.

  1. Create a blueprint: translate project requirements into a formal, machine-readable schema.
  2. Define the data model: apply a consistent architectural pattern to generate data models from that schema.
  3. Implement the endpoint: use the data models to create API endpoints.
  4. Implement the test: generate property-based tests to check the endpoint is correct and robust.
  5. Iterate and expand: re-apply the established patterns to build out the rest of the application.

Let's walk through each one.

Step 1: create a blueprint from the key ideas

Every project starts with ideas scattered across meeting notes, design documents and whiteboard sketches. The first step is to consolidate them into a formal database schema, which becomes the foundational memory document for every step after it.

  • Goal: translate project requirements into a formal database schema.
  • Pattern: a general LLM capability, generating an entity-relationship diagram (ERD) in Mermaid markdown from unstructured text.
  • Context: our project notes about the required entities and their relationships.
  • Example prompt: "Given these notes about benefit types, applications, and eligibility criteria, generate a Mermaid ERD for a database schema."

The outcome is a database_schema.md file that serves as a single source of truth for our data structures.

Example blueprint: the schema

erDiagram
    BENEFIT_TYPES {
        UUID id PK
        STRING name
        TEXT description
    }

    APPLICATIONS {
        UUID id PK
        UUID benefit_type_id FK
        ENUM status
        TIMESTAMP submitted_at
    }

    ELIGIBILITY_CRITERIA {
        UUID id PK
        UUID benefit_type_id FK
        STRING criterion
    }

    APPLICATION_DATA {
        UUID id PK
        UUID application_id FK
        STRING data_type
        STRING value
    }

    BENEFIT_TYPES ||--o{ APPLICATIONS : "are for"
    BENEFIT_TYPES ||--o{ ELIGIBILITY_CRITERIA : "have"
    APPLICATIONS ||--o{ APPLICATION_DATA : "contain"

Step 2: define the data model with the command-query model pattern

With a blueprint in hand, the next step is to implement the database models. To keep them consistent we use a predefined architectural pattern for our SQLModel classes.

  • Goal: implement a single, high-quality database model.
  • Pattern: the command-query model pattern (Base, Create, Table, Update), an architectural pattern we defined for SQLModel.
  • Context: a table definition from our database_schema.md.
  • Example prompt: "Using the 'EligibilityCriteria' table from @database_schema.md and our documented Command-Query model pattern, generate the corresponding SQLModel classes in @models.py."

For this to work the LLM needs both the database schema and a clear explanation of our model pattern. The pattern is heavily inspired by the official SQLModel tutorial on multiple models with FastAPI, which is itself an excellent document to hand the LLM as context.

Example command-query model

The LLM will often produce code that is correct but verbose. Here it puts fields such as id, created_at and updated_at directly into each model.

The LLM's first pass

# LLM's first pass is functional, but repetitive.
class EligibilityCriterionBase(SQLModel):
    criterion: str
    benefit_type_id: uuid.UUID = Field(foreign_key="benefit_types.id")

class EligibilityCriterionCreate(EligibilityCriterionBase):
    pass

class EligibilityCriterion(EligibilityCriterionCreate, table=True):
    __tablename__: str = "eligibility_criteria"
    id: uuid.UUID = Field(default_factory=uuid.uuid4, primary_key=True, index=True, nullable=False)
    created_at: datetime | None = Field(default_factory=lambda: datetime.now(UTC), nullable=False)
    updated_at: datetime | None = Field(default_factory=lambda: datetime.now(UTC), nullable=False)


"""For a PATCH request, all fields should be optional."""
class EligibilityCriterionUpdate(SQLModel):
    criterion: str | None = None
    benefit_type_id: uuid.UUID | None = None

That works, but it repeats itself. Change how IDs or timestamps are handled and you have to edit every model. Extracting the common fields into a BaseRecord model is better. You can instruct the LLM to do it, or do it yourself.

Human-refined code

Here is the pattern applied to an EligibilityCriteria model. It separates the base fields, the creation schema, the database table model and the update schema.

class BaseRecord(SQLModel):
    id: uuid.UUID = Field(default_factory=uuid.uuid4, primary_key=True, index=True, nullable=False)
    created_at: datetime | None = Field(default_factory=lambda: datetime.now(UTC), nullable=False)
    updated_at: datetime | None = Field(default_factory=lambda: datetime.now(UTC), nullable=False)

class EligibilityCriterionBase(SQLModel):
    criterion: str
    benefit_type_id: uuid.UUID = Field(foreign_key="benefit_types.id")

class EligibilityCriterionCreate(EligibilityCriterionBase):
    pass

class EligibilityCriterion(EligibilityCriterionCreate, BaseRecord, table=True):
    __tablename__: str = "eligibility_criteria"

"""For a PATCH request, all fields should be optional."""
class EligibilityCriterionUpdate(SQLModel):
    criterion: str | None = None

Step 3: building the endpoints

With the data models in place we can create the API endpoints. This step follows a standard pattern too.

  • Goal: create generic CRUD functions and a specific API endpoint for a model.
  • Pattern: a standard FastAPI router with GET, POST, PATCH and DELETE endpoints.
  • Context: our command-query EligibilityCriterion models from models.py.
  • Example prompt: "Implement standard RESTful CRUD operations in a FastAPI router for the 'EligibilityCriterion' table. Use the appropriate Schemas from @models.py."

The SQLModel tutorial on integrating with FastAPI gives a complete example of connecting these models to CRUD endpoints. It does need some setup first, such as a database dependency for FastAPI. We recommend an in-memory SQLite database during development, for speed and simplicity.

Example CRUD router: from specific to generic

The first output defines the endpoints but leaves the implementation out. Asked to fill them in, an LLM will generate specific, verbose logic for each function.

The LLM's first pass: specific CRUD logic

@criterion_router.post("/")
async def create_criterion(
    session: AsyncSessionDep, criterion: EligibilityCriterionCreate
) -> EligibilityCriterion:
    """Creates a new criterion. This is verbose and will be repeated for every model."""
    db_criterion = EligibilityCriterion.model_validate(criterion)
    session.add(db_criterion)
    await session.commit()
    await session.refresh(db_criterion)
    return db_criterion

# ... (and imagine similar verbose implementations for read, update, and delete)

This repeats itself heavily. Generic functions that handle CRUD for any SQLModel table are much cleaner, and the human programmer is the one to write them.

Human-refined code: generic CRUD functions

T = TypeVar("T", bound=SQLModel)

async def generic_create(schema: type[T], data: BaseModel, session: AsyncSession) -> T:
    new_data = data.model_dump(exclude_unset=True)
    try:
        insert = schema.model_validate(new_data)
    except ValidationError as e:
        raise HTTPException(status_code=422, detail=str(e)) from e

    session.add(insert)
    await session.commit()
    await session.refresh(insert)
    return insert

# ... imagine similar generics for Read, Update, Delete ...

With those in place the router becomes far simpler to read and maintain.

The refactored router

criterion_router = APIRouter(prefix="/criteria", tags=["criteria"])

@criterion_router.post("/")
async def create_criterion(
    session: AsyncSessionDep, criterion: EligibilityCriterionCreate
) -> EligibilityCriterion:
    return await generic_create(EligibilityCriterion, criterion, session)

@criterion_router.get("/")
async def read_criteria(
    session: AsyncSessionDep, skip: int = 0, limit: int = 100
) -> list[EligibilityCriterion]:
    return await generic_get_all(EligibilityCriterion, session, skip, limit)

@criterion_router.get("/{criterion_id}")
async def read_criterion(
    session: AsyncSessionDep, criterion_id: UUID
) -> EligibilityCriterion:
    return await generic_get(EligibilityCriterion, criterion_id, session)

@criterion_router.patch("/{criterion_id}")
async def update_criterion(
    session: AsyncSessionDep,
    criterion_id: UUID,
    criterion_update: EligibilityCriterionUpdate,
) -> EligibilityCriterion:
    return await generic_update(EligibilityCriterion, criterion_id, criterion_update, session)

@criterion_router.delete("/{criterion_id}")
async def delete_criterion(
    session: AsyncSessionDep, criterion_id: UUID
) -> EligibilityCriterion:
    return await generic_delete(EligibilityCriterion, criterion_id, session)

Step 4: implement high-fidelity, property-based tests

A robust API needs robust tests. We use property-based testing to check our endpoints hold up across a wide range of inputs.

  • Goal: make sure the API is robust, reliable and correct.
  • Pattern: property-based testing with pytest and Hypothesis.
  • Context: the create-criterion endpoint, its EligibilityCriterionCreate schema, examples of using httpx.AsyncClient and examples of Hypothesis.
  • Example prompt: "Write tests for the criterion router as defined in @crud.py. Use Hypothesis to generate test data based on the Schemas defined in @models.py and use the async test_client to call the API."

This step relies on testing infrastructure being in place, such as a pytest fixture for the httpx.AsyncClient used as a test client.

Example property-based testing

"""Generate Objects to Match the `Create` Model"""
@st.composite
def criterion_strategy(draw: st.DrawFn) -> dict[str, Any]:
    """Generate valid EligibilityCriterion data."""
    return {
        "criterion": draw(st.text(min_size=1, max_size=200)),
        "benefit_type_id": uuid.uuid4(),
    }
    """
    In a production test suite, we'd replace uuid.uuid4()
    with a pytest fixture that creates a BENEFIT_TYPES record
    and provides its ID, ensuring our foreign key constraint
    is always satisfied during testing.
    """

"""Supply the Object factory to the test to quickly test all properties"""
@given(new_criterion=criterion_strategy())
@pytest.mark.asyncio
async def test_create_criterion(
    self, new_criterion: dict[str, Any], test_client: AsyncClient
) -> None:
    """Test creating an EligibilityCriterion."""
    # In a real test, you'd ensure the benefit_type_id exists.
    validated_input = models.EligibilityCriterionCreate.model_validate(new_criterion)
    response = await test_client.post(
        "/criteria/", content=validated_input.model_dump_json()
    )

    assert response.status_code == HTTPStatus.OK
    created = response.json()

    assert "id" in created
    assert "created_at" in created
    assert "updated_at" in created
    assert created["criterion"] == new_criterion["criterion"]
    assert created["benefit_type_id"] == str(new_criterion["benefit_type_id"])

Example test client fixture

import pytest_asyncio
from httpx import AsyncClient, ASGITransport

from main import app, get_db

# ... other fixtures or test DB setup etc. ...

@pytest_asyncio.fixture
async def test_client() -> AsyncGenerator[AsyncClient, None]:
    app.dependency_overrides[get_db] = get_test_db
    transport = ASGITransport(app)
    async with AsyncClient(transport=transport, base_url="http://test") as client:
        yield client
    app.dependency_overrides.clear()

Step 5: iterate and expand

This is where the work pays off. With the patterns established by one high-quality implementation, scaling up is quick. We ask the LLM to apply those same patterns to the other tables in our schema.

Prompt 1: create models

  • Pattern: our command-query model pattern.
  • Context: all other models in @database_schema.md, using @models.py as an example.
  • Prompt: "Based on the examples in @models.py, implement all other models as defined in @database_schema.md."

Prompt 2: create CRUD logic

  • Pattern: our FastAPI router pattern.
  • Context: all other models in @models.py, using @crud.py as an example.
  • Prompt: "Based on the examples in @crud.py, implement CRUD logic for all models defined in @models.py."

Prompt 3: create tests

  • Pattern: our property-based testing strategy.
  • Context: all new endpoints in @crud.py, using @test_crud.py as an example.
  • Prompt: "Based on the examples in @test_crud.py, implement tests for all endpoints defined in @crud.py."

A few well-crafted prompts generate a large part of the application's boilerplate, which leaves more time for the logic that actually needs thinking about.

Guiding principles for LLM co-creation

A few principles hold this workflow together.

Design code, don't write it

Your role shifts from writing code to designing it. By establishing one high-quality example of a pattern, first one and then many, you give the LLM a template to follow. Your job becomes reviewing, refining and guiding, rather than typing out boilerplate, and you spend more time on architecture.

Control context and use sources

LLMs do their best work with small, focused contexts. Start each step with a clean context window. Use external documents, such as our database_schema.md, as a persistent memory you can feed back in. When working with a library, give the LLM its official documentation and tutorials: for the SQLModel classes and FastAPI endpoints we hand it the official tutorial on multiple models. And do not hesitate to restart a chat when the LLM gets sidetracked.

Let the AI clean its own mess

Use static analysis and automated testing. An LLM can produce code that looks right and fails under scrutiny.

  • Don't let the LLM repeat itself. When you see repetitive code, stop and work with it to create a generic function or a factory instead.
  • Tell it to verify its own work. Adding "Verify your work by running make ci" to a prompt does a lot.

Here is the Makefile we use for our static analysis suite:

.PHONY: static_analysis
static_analysis:
 @uvx ruff format .
 @uvx ruff check . --fix
 @uvx complexipy --details low src/app tests/
 @uv run --all-groups --with pip-audit pip-audit -l
 @uv run --all-groups --with pyright pyright src/app tests/

.PHONY: test
test:
 @uv run --group testing pytest -n auto -m "not slow" --cov=src/app --cov-report=xml
 @uvx diff-cover coverage.xml --fail-under=80 --compare-branch=main

ci: | static_analysis test

It runs everything at once with make ci, and static analysis on its own with make static_analysis, which is useful while the tests are not written yet. We use uv and its tool runner uvx, but equivalent tools work just as well.

Conclusion

This five-step workflow changes the development process. By acting as architects who define the patterns and guide the LLM, we automate the generation of high-quality boilerplate and keep our attention on the complex, unique parts of an application.

A structured approach like this turns the LLM from a code completion tool into a genuine development partner, and closes some of the gap between what these models can do and what real software engineering demands.

Colin Kerkhof

Author

Colin Kerkhof

I am an AI Engineer and I build production-ready AI systems.

LinkedIn

Keep up with the latest

Sign up for our newsletter and get our views on the latest in data & AI.

Let's talk about your data.

Get in touch with our team at contact@mozaik.ai or use the form below.

Or visit us at our office: Pakhuis De Hoop, Breestraat 59, Amersfoort.

Pakhuis De Hoop, Breestraat 59, Amersfoort

We only use your details to reply to your message. Read more in our privacy policy.