Institutional grade AI and automation for banks and credit unions that cannot hire an AI team — delivered as a service, and governed like a regulated estate.
Initial capability pack · August 2026
A community bank, credit union or building society is held to the standards of an institution many times its size, by a team that has no spare Tuesday.
Compliance and reporting obligations expand every year. Net interest margin does not, so the cost comes out of hours you do not have.
Large banks and digital challengers release automated features on a monthly cadence. The gap is not ambition; it is engineering capacity.
Specialist AI engineers are priced for large technology firms and leave for them. One hire is a single point of failure; a team is out of reach.
Staff paste customer detail into public chatbots today, because the tool is free and the deadline is real. You already own that risk.
The capability can be added without the headcount.
Every AI output is evidence for a human decision — never the decision itself.
Everything downstream inherits the quality of what is upstream. Governed data makes trustworthy knowledge; trustworthy knowledge makes intelligence you can act on; evidence you can trace makes decisions you can defend to an examiner — and assurance stops being a project.
Each capability slide highlights the stage it serves, and the diagram is navigable: click a stage to jump to it. The point is not any single feature — it is that the stages share one foundation, so every improvement compounds into the next.
Plenty of vendors sell a feature. We build the connected capability — one foundation, one governance story, and one conversation with your auditor instead of nine.
All of it runs on local models at zero per token cost where the work is sensitive, with measured throughput and instrumented spend. Appendix A3
Everything above is drawn from our own working estate — built and operated by the team that would build yours. No client of ours is running any of it; there are no case studies here but ours. The gold chips link to the appendix slide that justifies each claim.
Sensitive workloads run on models on hardware you own. The interesting part is not the privacy claim — it is that once the model is local, three separate procurement arguments stop existing.
Scoring, classification and drafting run at zero marginal cost, so volume is an engineering question rather than a budget one.
The data never leaves your infrastructure, so there is no cross border transfer to assess and no processor to name.
No vendor holds your prompts. Nothing is trained on and nothing is logged elsewhere, so nothing has to be trusted.
Hosted models are used only for non sensitive work, under a written policy. The backend is chosen per machine from a central registry.
Our whole estate runs this way daily: local inference for anything private, measured throughput, and spend instrumented per run. Appendix A3
The catalog is generated from the live database, not written by hand. Documentation that disagrees with reality is a defect the generator cannot produce.
An automated schema standard checker runs over the estate and fails on violations rather than warning about them. It has caught real defects, not hypothetical ones. Appendix A2
Collectors may write only raw areas. Readers cannot write at all. This is enforced by the database engine, not by convention or code review.
Collected data is never edited in place; issues are flagged alongside it. Every structural change is a numbered, replayable migration.
Evidence: our own estate — 700M+ rows, ~200 GB, 44 registered stores under exactly this discipline. Appendix A1
A benchmarked multi engine pipeline reads scans, PDFs, spreadsheets and handwriting. The engine used is recorded per document, so a bad read is traceable rather than mysterious.
A knowledge graph and a semantic index over the whole corpus, so "what do we know about X?" returns a cited answer instead of a folder path.
Benchmark question sets with known answers are scored before anyone relies on the system, and rescored on every change. Trust is tested, not assumed.
Proven on genuinely private material: an 18,000 file email and notes archive indexed into 68,464 retrievable chunks, with people profiles, organisation profiles and thread reconstruction — processed on premises. Appendix A4 Appendix A5
The machinery is demonstrated on investment watchlists in our own estate. Pointing it at a bank's world is configuration and domain work, not new invention — and we will say so on the day rather than after the invoice. Appendix A10
The community bank reality: a core system with no modern interface, vendor portals that export nothing useful, and green screens that a person retypes from. Waiting for the integration is a multi year answer to a Tuesday problem.
A vision model reads a broker screen and normalises what it sees into structured data. No integration project, no vendor negotiation, no export button required.
The same pattern reads any screen or scanned form: a terminal, a portal, a statement, a form a member filled in with a pen.
A bridge while proper integrations are negotiated — not a replacement for them. Screens change, and a screen reader is a maintenance commitment. We say so before you buy it.
The honest framing matters here: this buys you the two years you would otherwise spend waiting on a vendor roadmap. Appendix A7
An agent that refuses ambiguous work. Faced with an instruction that could mean two things, it asks one specific question rather than guessing confidently and producing something plausible and wrong.
The email agent is structurally prevented from sending as a person. Allow and block lists are validated at startup, and an overlap between them is a hard error that stops the service rather than a warning nobody reads.
Low confidence paths escalate to a review queue instead of proceeding. Read only access wherever writing is not required, so a compromised component cannot corrupt anything.
In a regulated shop, what the AI refuses to do matters more than what it can do.
These are properties of the system, not instructions in a prompt. A guard that can be argued out of its position by a cleverly worded input is not a guard.
This is the least controversial place to start and the easiest to measure: no credit decision is touched, and the hours are visible in the first fortnight.
Every scheduled job and service on one health panel, with freshness alarms. The absence of a run is itself checked, because a silent gap is the failure nobody notices.
Twice daily, verified, with a nightly off site copy and a ledger that alerts when a link in the chain goes missing. Appendix B4
Every AI answer cites its sources; every decision chain is logged; every report can be regenerated from the data it claims to describe. Appendix A2
We found our own sync tool silently resurrecting deleted files, and 47% of one store redundant. So we rebuilt the strategy around a rule: sync is not backup.
We lead with that because it is the honest origin of the discipline. A supplier with no failure story has either not run anything for long, or is not telling you about it. Appendix B4
| Value lever | Where it lands in a community bank | Illustrative before → after |
|---|---|---|
| Cut cost | Document intake, meeting minutes, reporting assembly, routine correspondence, file retrieval | Committee minutes and actions: ~50 minutes → ~3 minutes. A recurring report: days of assembly → regenerated from live data. |
| Scale capacity | Monitoring coverage, document volume per person, response time to members and customers | Watching what matters: a few names by hand → always on, with thresholds. Document throughput: rises without new hires. |
| Raise quality | Provenance on every number, institutional memory that survives turnover, fewer manual transcription errors | "Where did this figure come from?" → answered with lineage. A leaver's knowledge: gone → still answerable. |
Deliberately qualitative. We will not put a saving on your balance sheet before we have seen your documents — the measurement is the first thing an engagement builds, not the first thing it claims.
No pricing appears in this pack. Commercial structure is shaped in conversation once a first use case is chosen.
| When | What happens |
|---|---|
| 07:30 | The overnight brief is waiting: two mentions of the bank worth reading, one vendor notice, one local employer story — each with a thirty second summary and a link to the source. |
| 09:00 | A scanned lending bundle lands as structured data. Fields extracted, engine recorded per page, and four exceptions flagged for a human rather than guessed at. |
| 11:00 | Someone asks: "what did we tell the examiner last year about our model governance?" — a cited answer in seconds, quoting the institution's own file. |
| 14:00 | Committee meeting. Minutes and owned action items are distributed before people are back at their desks. |
| 16:30 | A recurring report is regenerated from live data, with every figure traceable to the table it came from. A person checks it and signs it. |
| 17:30 | The team goes home. Monitors keep watching, backups run and verify, and the health panel is green in the morning or it says why. |
Nobody's job changed. Everybody's reach did.
A scoped four week embedded engagement: two of our consultants alongside your frontline team, on a use case chosen together in week one. Training and build are the same activity, so the capability stays when we leave.
Your AI function, without the headcount. A weekly working session, a visible request board, and everything documented inside your own estate. Small things ship fast; big things get scoped honestly. Appendix B2
The meeting pipeline, document intake and monitoring, stood up as services on your infrastructure. The fastest visible win, and the least disruption.
Most AI projects fail for organisational reasons rather than technical ones — published failure rates run between 80% and 95%. Sponsorship decays, nobody agreed what success was, the data was not ready, the pilot never scaled. Our engagement is designed explicitly against those documented failure modes, which is why every door starts with a measurement and a gate. Appendix B1
Zero trust access for anything browsable: allowlisted addresses, one time PIN, identity checked on every request. Nothing internal is exposed to the open internet.
Dashboards and viewers hold no credential that could write, so a compromised page cannot corrupt data. Writing is a separate, deliberately narrower path.
The database engine itself refuses the write, rather than the application choosing not to attempt it. Collectors write only raw areas; readers cannot write at all.
Sensitive material is processed by local models on your hardware. Cloud AI touches only non sensitive workloads, and only under a policy agreed in writing in advance.
Raw data is never edited; every structural change is a versioned, logged migration; twice daily verified backups with a nightly off site copy and a ledger that alerts if the chain breaks.
This is how our own estate already runs. Governance was not added for this deck. Appendix B3
Not a proposal to integrate other people's products — a working estate: governed data, a knowledge layer, monitors, document pipelines and guarded agents, operating together today. This pack is drawn from it.
Published negative results, real defects caught by our own audits, and written post mortems on what went wrong. You will always know what the system cannot do, which is what makes the rest believable.
Community banks and credit unions rarely sustain a full time AI engineering function. We supply it as a service and document everything inside your estate, so it outlasts any individual — ours or yours.
We build your capability alongside your people: documented, catalogued and transferable. An asset the institution owns, not a black box it rents.
A working session with the people who do the work: where do the hours and the risks actually live, and which of them would you most like back?
Consulting sprint, embedded retainer, or a productised tool. Whichever one fits the answer to question one. Appendix B2
A first working capability inside 2 to 4 weeks of the build starting, with a go or no go gate at every stage. You are never committed beyond the last proven step. Appendix B1
Initial capability pack for discussion — not a formal engagement proposal. Prepared in the United Kingdom. ApisQ provides technology and workflow capability; it is not a regulated financial services provider, and nothing here is compliance, legal or investment advice. All examples are drawn from our own working estate.
For the person who has to defend this internally. Everything here describes our own estate; where a banking application is not yet built, it is tagged In build or Roadmap and says what would be new.
A1 to A10 — how each piece of machinery actually works, the numbers behind it, and an honest mapping onto banking use cases
B1 to B5 — phasing and gates, the four week engagement, build versus buy, resilience, and what an engagement needs from you
C1 to C4 — the questions we expect from a board and a risk committee, plus plain language definitions
The catalog is produced from the live database, and a staleness check fails the build when the generated file and the database disagree. Hand written data documentation is always out of date; the only question is whether anyone has noticed yet.
It is also machine readable, which is what lets an AI assistant discover the right table instead of being told about it in a prompt that will rot.
Exactly as collected. Immutable — only collectors may write, nobody may edit
Cleaned and derived, always rebuildable from raw, so corrections never destroy the original
Expensive derived knowledge — graphs, indexes, reports — backed up accordingly
Dashboards and answers, each number traceable back left through the tiers
A check that has never failed is not evidence of health; it is an untested check. Ours have failed twice, on real problems, in production.
A collector module was creating its own tables at runtime, outside the migration system — so the database's structure could change without a reviewed, numbered, replayable migration.
The schema standard checker failed on it. Table ownership moved to migrations, and the rule is now enforced rather than documented.
Redundant indexes covering the same columns, costing 1.9 GB and slowing every write to the affected tables. Invisible to anyone reading the code; obvious to a check that reads the live catalog.
A violation exits non zero and stops the pipeline. A warning is a defect with a longer lifespan.
Collectors write raw areas only; a read role cannot write anywhere. The database refuses, so no application bug can bypass it.
Numbered, replayable, reviewed, with rollback scripts beside them — because a change you cannot undo is a change you cannot safely make.
The point is not that local models are always better. It is that the sensitivity decision is made once, in writing, and then enforced by the plumbing rather than remembered by whoever is on shift.
A routing table sends each document to the engine that benchmarked best for its type: born digital PDFs, scanned pages, spreadsheets, presentations, handwriting. Where two engines disagree materially, the document is flagged rather than silently resolved.
The engine used is written into the output filename, so a quality complaint six months later is a lookup rather than an investigation.
Two A/B comparisons published internally, with the losing engine's results kept. The post mortem documents the failure modes we hit — multi column layouts, faint scans, tables that survive as text but not as structure — so nobody rediscovers them at a client's expense.
Accuracy is a measured property of a pipeline on your documents, never a vendor's headline number.
The system assembles a profile for each person and organisation that appears across the corpus, resolving name variants, so a question about a counterparty returns one answer rather than five partial ones.
Fragmented correspondence is rebuilt into conversations in order, which is what turns an archive from a search box into an account of what actually happened.
A benchmark evaluation set with known answers is scored before the base is used, and rescored on every change. Retrieval quality is a number we can show you, not a claim.
This is the closest analogue in our estate to a bank's own institutional memory: private, unstructured, decades deep, and the thing everyone assumes is unsearchable. Nothing in it left the machine it was processed on.
Scheduled collectors pull from news, public attention and social sources into the raw tier
A language model reads each item for tone toward the target and for substance, not keyword hits
Only a meaningful move raises anything. The bar is configured per target and reviewed
An hourly narrative summarising what changed, readable in thirty seconds
Email and SMS or messaging channels, with the evidence attached
Six monitoring services in scheduled production in our estate, running on a timer with health reported to the ops panel.
Six independent public news and attention sources at no subscription cost. The expense in this layer was always the engineering, never the data.
An alerting system that cries wolf is worse than none, because it trains people to ignore the channel you will one day need them to read.
Social sources are read through a licensed social data provider under its terms. Sentiment is scored by local models, so the content of what is read never leaves the estate.
A screenshot of a broker terminal goes in; structured, normalised position data comes out. The model reads the screen the way a person does, and the output is validated against expected shapes before anything downstream trusts it.
No integration project, no API, no vendor negotiation, no export button. It was standing up in days, not quarters.
Generalises to any screen or scanned form: a terminal, a vendor portal, a statement, a paper form. The pattern is the asset; the broker screen is just where we proved it.
Recording collected from the meeting platform, with speaker separation
Local transcription with speaker labels merged back onto the text
Below the bar, the pipeline stops. A poor transcript never becomes a confident summary
Decisions, discussion and actions separated, each traceable to a passage
Owners mapped, duplicates merged, tasks published with due dates
The quality gate is the part worth copying. Every stage that follows it assumes a usable transcript, so allowing a bad one through does not produce a slightly worse summary — it produces a confident, wrong one that somebody actions.
Every message classified into agreed categories, with a confidence score. Low confidence goes to a person rather than to a guess.
Actionable requests become tasks with an owner and a due date; categories route to the channel that owns them; noise is suppressed rather than forwarded.
The agent cannot send as a person. Allow and block lists are validated at startup and an overlap between them is a hard error that refuses to start the service.
Mailbox access uses standard delegated authorisation with the narrowest scope that works, refreshed automatically and revocable by your administrator without involving us. Credentials live in a managed vault, never in code or configuration files.
A prompt instructing a model not to send mail is a suggestion. A service that will not start when its own configuration is contradictory is a control. In a regulated institution only the second one is worth writing down.
This is where the banking use cases live. None of them is Demonstrated, because no bank is running them. The machinery underneath each one is.
| Application | Demonstrated machinery it reuses | What would be new | Maturity |
|---|---|---|---|
| Document intake for lending files | Multi engine transcription, extraction, quality gates | Lending document schemas; exception rules agreed with credit | In build |
| Onboarding and KYB document collection | Document pipeline, entity extraction, knowledge graph | Registry connectors; a verification policy you own | Roadmap |
| Covenant and early warning monitoring | Collectors, model scoring, thresholds, alerting | A covenant model per facility; a borrower data feed | Roadmap |
| Member and customer communications | Grounded drafting with citations; the send authority guard | Tone and disclosure rules; channel integration | Roadmap |
| Complaints triage | Classification, task extraction, routing, follow up radar | A complaints taxonomy; regulated response clocks | Roadmap |
| Regulatory report assembly | Governed warehouse, lineage, regeneration from live data | Report templates; a sign off workflow with named owners | Roadmap |
Demonstrated today and reusable by every row above: the same source material is regenerated in 3 output languages (English, 中文, العربية) without rewriting the analysis underneath.
Map the workflows, inventory the documents, pick the first use case together
First working outcome, visible to the team that will use it
Extend along the spine, each capability reusing the last one's foundation
Your people run it; the capability and its documentation are yours
Discover exists to establish a baseline. Without one, no later claim about hours saved can be defended to a board.
The first capability is chosen to be visible to the people doing the work within a fortnight, because sponsorship decays fastest in the gap before anything appears.
Changes pass an agreed evaluation before adoption, with kill criteria written down in advance.
Two audiences, one programme. Senior stakeholders get concepts, governance and demonstrations; the frontline team builds a real MVP with two of our consultants beside them.
Choose the use case together and baseline the current process. Senior session on what AI is, is not, and what governance it needs.
Build with the frontline team on real material. First rough output in front of the people who will judge it.
Harden: guards, exceptions, the review queue, the measurement. Second senior touchpoint, on evidence rather than slides.
Working MVP handed over and documented in your estate, with a written go or no go recommendation.
Every element above answers a documented failure mode. Sponsorship decay: a visible outcome inside four weeks. No measurement: the baseline is taken in week 1. Data unreadiness: a use case needing data you lack is rejected early. The scaling wall: the MVP sits on the foundation production would use.
A course teaches concepts that fade. Building the thing with your people, on your material, leaves them a working system and the ability to change it — documented in your own estate as it is built, so the knowledge does not live in whoever was in the room.
Commodities: public data feeds, foundation models, meeting capture plumbing, cloud infrastructure. There is no pride of authorship where a vendor is genuinely better value, and rebuilding a commodity is how budgets disappear.
The connective tissue: your data foundation, your knowledge layer, monitors pointed at what you care about, and the guards. These are the parts no vendor can sell you, because they encode how your institution works.
The integration layer and the catalogs are documented and handed over. That is the defence against vendor sprawl, and it is the difference between an asset and a subscription.
Any paid vendor is admitted only after a written evaluation naming what was tested, against what alternative, and what would make us change our mind. The evaluation goes in your estate, so a future procurement question has an answer rather than a memory.
Fewer vendors is itself a control: fewer data sharing agreements, fewer processors, fewer audit surfaces, fewer places a breach can start. Consolidation is a governance argument before it is a cost one.
Our file sync tool was silently resurrecting deleted files: a deletion on one machine reappeared from another, so the estate never actually shrank. An audit found 47% of one store was redundant.
A sync mirrors your mistakes with admirable speed. A backup keeps a version from before you made them. We rebuilt the strategy around that distinction, and the database is now covered by the dump and off site chain specifically, not by the file sync.
We include this because it is the most useful thing we know about resilience, and we only know it because we got it wrong first.
One person who owns the engagement and joins the weekly session. Roughly half a day a week, most of it deciding rather than doing.
Read only access to the systems in scope for the chosen use case, granted a step at a time and never all at once. Revocable by you, without us.
A representative sample of the real material, so quality is measured on your documents from day one rather than on a demonstration corpus.
Written down in week one: what, observed by whom, would make this a clear win — and what result would make stopping the right answer.
Notably absent from that list: a data migration, a core system change, a committee, or a signed multi year commitment. If the first use case needs any of those, we have chosen the wrong first use case.
No. It changes what their hours buy. The system does the collecting, reading, transcribing and formatting; your people do the judging, the deciding and the conversations. The scarce resource in a community bank was never willingness — it was attention.
Trust is engineered, not assumed. Answers cite retrievable sources; knowledge bases ship with benchmark question sets scored on every change; claims pass adversarial review; and nothing reaches a member, a customer or a regulator without human sign off.
We build audit trails and explainability in from day one, so the evidence is easy to produce rather than assembled in a panic. You remain accountable, and we are not your compliance adviser — nothing in this pack is compliance or legal advice.
The third answer is the one we would rather over state than under state: a technology supplier cannot carry your regulatory obligation, and any supplier who implies otherwise is selling you a problem.
We would rather answer these two questions before the capability questions. In our experience they are what a board actually asks, and a supplier who is vague about either has not thought about your position.
| Term | What it actually means |
|---|---|
| LLM (large language model) | AI trained to read and write text. Useful as a tireless junior who has read everything, and who must always show its sources. |
| Agent | An AI given tools and a goal that decides its own steps — "find every file mentioning this vendor and summarise what changed" rather than one question, one answer. |
| RAG and GraphRAG | Retrieval augmented generation: the AI answers only from documents retrieved for that question, and cites them. GraphRAG adds a map of entities and relationships, so connected facts surface together. |
| Knowledge graph | A map of who and what your documents mention and how they connect — supplies, owns, acquired, complained. It answers questions a keyword search cannot. |
| Grounding and citations | Forcing the AI to answer from retrieved source text and to show where each claim came from. The single most important control on trustworthiness. |
| Hallucination | An AI stating something fluent and false. Grounding, citations and benchmark testing manage the risk down to something auditable. |
| Local model and on premises inference | An AI model running on hardware you own. The data never leaves the building and there is no per use fee. |
| Vision model | An AI that reads images and screens rather than text files — how a green screen or a scanned page with no export button gets ingested. |
| MCP (model context protocol) | A standard way to expose a tool or a data source to an AI assistant, so capability is added by connecting a defined interface rather than by pasting instructions into a prompt. |
| Term | What it actually means |
|---|---|
| Least privilege | Each system and each person gets the minimum rights the job needs, so a read only tool cannot write even if it is compromised. Enforced by the database server, not by an application's good intentions. |
| Immutable history and flag, do not fix | Collected data is never edited in place. Problems are flagged alongside it and corrections applied in a separate layer, so nothing is silently rewritten. |
| Point in time data | Data as it was known on the day, including figures later restated. Without it, any analysis of the past quietly uses knowledge nobody had at the time. |
| Audit trail, lineage and provenance | The recorded path from a number on a screen back to the source document it came from, and the log of who or what changed it. It turns "where did this come from?" into a lookup. |
| Zero trust access | Every request verifies identity first; nothing is trusted merely for being inside the network. It puts the boundary at the edge rather than inside each application. |
| Shadow AI | Staff using unsanctioned public AI tools for work, usually because the sanctioned option is missing or slower. The durable fix is a better sanctioned tool, not a stricter memo. |
| Human in the loop | A design where the system produces evidence and a person makes the decision. In its stronger form, low confidence cases escalate automatically rather than waiting to be noticed. |
We would welcome the conversation about where the hours are going, and which of them you would like back first.
info@apisq.co
Initial capability pack, v1 · August 2026 · Prepared in the United Kingdom for discussion purposes only. This document is not a formal engagement proposal. ApisQ provides technology and workflow capability; it is not a regulated financial services provider, and nothing here constitutes compliance, legal or investment advice, research, or a financial promotion. Capability claims tagged Demonstrated describe our own working estate; claims tagged In build or Roadmap are not yet delivered. No client implementation is described or implied anywhere in this pack.
Read the main pack; dip into the appendix only where you want the evidence. Most readers spend about fifteen minutes on the first twenty slides and never need the rest — but every claim in them links to the appendix slide that backs it up.
Scrolling moves one slide at a time, so nothing lands halfway. Arrow keys, space and page keys do the same, as do the ▲ ▼ buttons at the bottom right. Click the gold page number to jump anywhere.
👍 / 👎 rates the slide you are on in one click; 💬 Feedback leaves a written note. Both genuinely shape the next version.
Any dotted gold term jumps to its plain-language definition in the glossary, and a ← Back button returns you to exactly where you were.
This pack is public and asks for no sign in, so we do not know who you are. It records anonymous usage analytics — how long each slide is open, plus any reaction or note you choose to leave. Nothing else is tracked, the timer pauses whenever you switch away, and it is used only to improve the next version.