Skip to content
Data preparation and records management

Engagement 01 · Data Preparation · Seven Weeks

Records your automated systems can actually depend on

When the data coming in is inconsistent, the output is unreliable regardless of how well the model was built. This engagement works on the records before anything else does.

Back to home

What this engagement delivers

A documented, cleaned dataset and the rules to keep it that way

For management

At the end of seven weeks, your organisation holds three things: a data dictionary written in plain terms that any department can read, a cleaned dataset in an open format, and a set of validation rules that catch problems in new records before they accumulate.

Staff no longer have to resolve the same formatting conflicts project after project. The groundwork is done once, properly, and documented so it holds.

For technical staff

Deliverables include a field-level data dictionary with agreed definitions, a deduplicated and formatted dataset delivered in CSV or Parquet, and rule specifications in a format compatible with your ingestion pipeline.

Validation logic is documented rather than locked into tooling, so your team can review, extend, and own it independently after handover.

The situation

Records gathered across years and systems rarely agree with each other

For management

Most organisations that have been running for more than a few years hold data in several systems that were never designed to work together. A customer record in one system means something different from a customer record in another. Fields are missing. The same entity appears under two slightly different names.

This is not a sign of carelessness. It is what happens when a business grows and acquires tools as it goes. But it does mean that any automated system reading those records is working from material that was never intended to be consistent.

For technical staff

Common issues include duplicate primary keys across migrated datasets, date fields formatted as strings, categorical fields with inconsistent casing, and reference columns pointing to IDs that no longer exist in the main table. Historical imports compound these problems because they were often done under time pressure without documentation.

Feature engineering built on this material introduces noise that cannot always be traced back to the source. Fixing it upstream is more reliable than compensating downstream.

The approach

Field by field, starting from what your systems actually contain

For management

The engagement starts by examining what records exist, which departments use them, and what each field is understood to mean in practice. Where definitions differ between teams, we work with the relevant staff to agree on a single version that reflects how the business actually operates.

Cleaning follows the agreed definitions. Nothing is changed without a written reason. The result is a dataset that can be traced back to decisions your organisation made, not arbitrary choices made during processing.

For technical staff

We profile each source table, document null rates, cardinality, and value distributions, then map conflicts across sources before writing any transformation. Deduplication is done using deterministic matching on agreed canonical identifiers rather than fuzzy methods that generate false positives.

Transformation scripts are written in SQL or Python depending on your stack, and are left with your team alongside the data dictionary. Validation rules are expressed as assertions that can be integrated into dbt, Great Expectations, or a bespoke CI check.

Working together

Seven weeks, structured around your team's existing schedule

Weeks 1–2

Inventory and scoping

We review source systems, sample data, and current documentation. A scoping note is written before any cleaning begins, naming what will and will not be addressed in this engagement.

Weeks 3–6

Definition, cleaning, validation

Field definitions are agreed with the relevant staff, cleaning runs against those definitions, and validation rules are drafted and tested on a sample before being applied to the full dataset.

Week 7

Handover and review

The data dictionary, cleaned dataset, and validation rules are reviewed together with your team. The session ensures the materials can be used and maintained without ongoing involvement from us.

Investment

¥40,000 JPY for seven weeks of work

What is included

Data dictionary

Field-level definitions agreed with your staff and written in plain terms. Updated throughout the engagement and handed over at close.

Cleaned dataset

Delivered in an open format (CSV or Parquet) with a log of every change made and the reason for each.

Validation rules

Documented assertions for keeping new records consistent, compatible with common validation frameworks and your existing ingestion process.

Handover session

A review meeting with your team at close to walk through the materials and confirm they can be used independently.

Practical details

Duration

Seven weeks from engagement start

Investment

¥40,000 JPY (invoiced at engagement start)

Format

Work sessions with your staff, remote or in Oita

Output ownership

All materials belong to your organisation at handover

Suitable for

Organisations whose records have accumulated across several systems and reorganisations

Measurement

What is measured before and after this engagement

Three figures are taken at intake and again at delivery against the cleaned dataset. They are stated in the scoping note and reviewed at handover.

Before

Duplicate record rate

Percentage of records appearing more than once under different identifiers. Measured per source table and across combined views.

After

Target: no duplicates remaining on agreed canonical identifiers

Before

Field completion rate

Percentage of required fields populated across all records. Reported per field and in aggregate, highlighting fields below an agreed threshold.

After

Target: fields below threshold documented with reason and action taken

Before

Definition conflicts

Number of fields with differing definitions across departments or source systems. Identified during the inventory phase and listed in the scoping note.

After

Target: each conflict resolved to a single documented definition

Our commitment

The scope is written before work begins and held to throughout

The scoping note produced at the start of this engagement names what is included and what is not. Neither side should be uncertain about what the engagement covers. If the agreed scope turns out to require adjustment as work progresses, that is discussed and documented before any change is made.

The initial conversation before any commitment is made carries no obligation. If after that conversation you decide this engagement is not the right fit for your situation, there is nothing to undo.

Before any commitment

Initial conversation

No charge, no obligation. We discuss your situation and whether this engagement is appropriate.

Written scope

A scoping note is produced and agreed before work begins. This is the reference point for both sides throughout.

Your ownership

All materials produced during the engagement belong to your organisation. No ongoing dependency on Naze Blend Zone is created.

Next steps

How to begin

Step one

Send a message

Use the contact form below, or write directly to info@naze-blendzone.com. Describe your situation in whatever terms feel natural — there is no required format.

Step two

Initial conversation

We will arrange a call to discuss your data situation, which systems are involved, and whether this engagement addresses what you are working with.

Step three

Scoping note and start

If the engagement is a fit, we produce a written scoping note for your review. Work begins once it is agreed. Seven weeks from that point.

Other engagements

Data preparation is one of three areas

Engagement 02

Model Monitoring Setup

Establishing ongoing monitoring for systems already deployed — accuracy drift, distribution shifts, alert thresholds. Five weeks. For organisations running systems introduced a year or more ago without continuing measurement.

¥36,000 JPY

View

Engagement 03

Governance and Policy Setup

Internal rules for automated tool use — approved applications, data handling restrictions, review requirements, record keeping. Five weeks. For organisations where staff have begun using tools ahead of any policy.

¥30,000 JPY

View

Data Preparation Engagement

Seven weeks to records your systems can rely on

If your organisation is dealing with inconsistent records across systems, this engagement addresses that directly. The initial conversation is without obligation.