Skip to content
Shrijayan
Go back

Trained a Transaction Foundation Model from Scratch on 2×RTX 4090

The Idea

NVIDIA published a blueprint for Transaction Foundation Models (TFM) — small transformer models trained on financial transaction sequences. They claimed a 36% lift in fraud detection by combining TFM embeddings with traditional features. I wanted to validate this end-to-end on consumer hardware.

The Setup

The Architecture

The model is a 29M-parameter Llama decoder-only transformer:

A decoder-only causal LM predicts the next transaction token given the history — forcing the model to learn transaction patterns, merchant relationships, spending cycles, and anomaly signals.

Architecture note: Decoder vs encoder matters. The TFM approach here is decoder-only (predicts next token autoregressively), but for pure representation learning like fraud, a BERT-style encoder (bidirectional attention) is often preferred — each transaction token sees both past and future context. NVIDIA chose decoder-only for generation capability alongside embeddings; your choice depends on whether you need sequence generation or just embeddings.

Training from Scratch

Instead of using NeMo (which had dependency conflicts), I built a standalone training script directly on HuggingFace transformers with PyTorch.

Training config:

Loss curve:

Step   000: loss 8.00
Step   200: loss 2.20
Step   400: loss 1.50  
Step   600: loss 1.20
Step   800: loss 1.05
Step  1000: loss 0.99  (val loss: 0.99)

Total time: 18.6 minutes. That’s it. A functioning transaction foundation model trained in under 20 minutes.

Extracting Embeddings

For each transaction, I constructed a sequence: [BOS] + tokenized_transaction + [EOS], fed it through the model, and took the last hidden state as a 512-dimensional embedding vector.

90,000 transactions processed in about 30 seconds.

Fraud Detection: Raw vs TFM vs Combined

I compared three XGBoost classifiers on the test set:

FeaturesROC-AUCAvg Precisionvs Raw
Raw 13d (baseline)0.99740.5451—
TFM Embeddings 64d0.97820.3616-33.66%
Combined (77d)0.99300.7205+32.18%

The TFM embeddings alone underperform raw features (expected — they capture broad transaction structure, not fraud-specific signals). But combined with raw features, they deliver a consistent +32% lift in average precision — closely matching NVIDIA’s published findings.

Why This Matters

  1. Consumer hardware is enough — 18 minutes on a single RTX 4090. Most teams already have this.
  2. TFMs complement existing models — you don’t replace your fraud model, you augment it.
  3. Domain-specific > generic LLMs — a 29M param transaction model beats any 7B general-purpose LLM on this task.
  4. End-to-end pipeline validated — from raw CSV → tokenization → training → embeddings → fraud classifier, all reproducible.

Key Takeaway

Small, domain-specific foundation models are incredibly practical. You don’t need a cluster or a foundation model API. One GPU, one evening, and you can train a transaction model that meaningfully improves fraud detection.

This validated that the approach works end-to-end on accessible hardware — and the +32% lift matches the direction of the published results.


Share this post on:

Previous Post
Vector Databases: A Survey of the Landscape
Next Post
Contributing Windows Support to Tailscale Browser Extension