Open-Sourcing Finsight: Adapting Qwen3-14B for Financial Reasoning on AWS
We are open-sourcing two finance-adapted checkpoints of Qwen3-14B (CPT and CPT-SFT). Here is our end-to-end training pipeline, 13-gram decontamination, benchmark trade-offs, and lessons learned.
Today, we are excited to open-source two specialized financial language model checkpoints developed at Maxint: Finsight-Qwen3-14B-CPT and Finsight-Qwen3-14B-CPT-SFT.
- Intermediate Checkpoint (CPT): maxint-inc/Finsight-Qwen3-14B-CPT
- Final Merged Checkpoint (CPT → SFT): maxint-inc/Finsight-Qwen3-14B-CPT-SFT
- Base Model: Qwen/Qwen3-14B-Base (Apache 2.0)
At Maxint, our goal with this research experiment was straightforward: test whether exposing an open base model to financial filings, news, and specialized instructions could improve financial reasoning within a compact, budget-conscious cloud run.
We executed a two-stage adaptation on AWS—combining Continued Pretraining (CPT) on raw financial corpora with Supervised Fine-Tuning (SFT) on curated finance instructions—all completed end-to-end within a single day.
In the spirit of open scientific inquiry, we are sharing not just the weights, but our full methodology, 13-gram test set decontamination pipeline, benchmark results, and an honest look at the trade-offs we encountered.
Executive Summary & Key Findings
- Strong Gains in Numerical Reasoning: On tabular financial calculation tasks, our adaptation achieved major leaps over the base model. FinQA accuracy surged by +563% relative (from 0.019 to 0.126, a +10.7 percentage point gain). Multi-turn financial reasoning (ConvFinQA) climbed from 0.148 to 0.203 (+37.2% relative gain).
- The Classification & Extraction Trade-off: While numerical reasoning leaped forward, unweighted average performance across six financial benchmarks dropped modestly from 0.441 to 0.399. The decline was heavily concentrated in aspect-based sentiment (FiQA-SA dropped from 0.681 to 0.370) and named entity recognition (NER F1 dropped from 0.238 to 0.180).
- SFT as a Stabilizer: Comparing checkpoints revealed that raw CPT severely disturbed the model’s classification calibration. Subsequent SFT successfully initiated a recovery across every single benchmark relative to the CPT-only stage, though it did not completely restore base classification baselines under raw zero-shot evaluation.
- Consumer GPU Friendly: Exported as 16-bit BF16 Safetensors (29.5 GB), the final model loads smoothly in 4-bit NF4 quantization on an inexpensive 16 GB GPU (such as an NVIDIA Tesla T4 on Google Colab) taking only 9.05 GB of VRAM.
Compute & Training Infrastructure
We ran the entire experiment on Amazon Web Services using a single p4d.24xlarge virtual machine configured as follows:
| Component | Specification |
|---|---|
| Instance Type | AWS p4d.24xlarge |
| Accelerators | 8 × NVIDIA A100 SXM4 (40 GB VRAM each) |
| Precision | BF16 mixed precision, gradient checkpointing |
| Attention | Flash-Attention 2 (with SDPA fallback) |
| Distributed Launcher | accelerate launch with single-node DDP (configs/accelerate_ddp.yaml) |
| VM Allocation | Approximately 1 day |
| End-to-End Elapsed Time | ~20 hours |
| Active Training Duration | CPT: ~5 hours · SFT: ~2 hours |
| Residual VM Time | ~13 hours (data loading, evaluation harness, checkpoint exports) |
The total active compute represented approximately 56 allocated GPU-hours for training, within an operational footprint of 160 allocated GPU-hours across the run.
The Data Pipeline
Financial language is dense, jargon-laden, and full of complex numerical relationships. To adapt Qwen3-14B-Base, we split the curriculum into two distinct stages: broad domain immersion, followed by structured task execution.
┌────────────────────────────────────────┐
│ Qwen3-14B-Base │
└───────────────────┬────────────────────┘
│
▼
Stage 1: Continued Pretraining (CPT)
• ~330M raw tokens (EDGAR, News, BAAI)
• QLoRA Rank 64 / Alpha 128
• 800 steps on packed 2,048 sequences
│
▼
┌────────────────────────────────────────┐
│ Finsight-Qwen3-14B-CPT │
│ (Intermediate Checkpoint) │
└───────────────────┬────────────────────┘
│
▼
Stage 2: Supervised Fine-Tuning (SFT)
• 191,428 instruction pairs (ChatML)
• QLoRA Rank 32 / Alpha 64
• 2 epochs with full token loss
│
▼
┌────────────────────────────────────────┐
│ Finsight-Qwen3-14B-CPT-SFT │
│ (Final Merged Checkpoint) │
└────────────────────────────────────────┘
1. Continued Pretraining Corpus (~330M Tokens)
Our raw financial text corpus was built using src/finsight/data/prepare_cpt.py by streaming from the training splits of three core sources, capped in configs/datasets.yaml:
| Dataset Source | Streamed Fields | Document Cap | Domain Contribution |
|---|---|---|---|
anonymous-md/EDGAR_FILINGS_DATASET | parsed_md | 7,000 | SEC 10-K annual reports, MD&A disclosures |
ashraq/financial-news-articles | title + text | 100,000 | Real-time market commentary and company news |
BAAI/IndustryCorpus_Finance | text / content | 80,000 | Long-form institutional finance and macroeconomic text |
Filtering & Processing Rules:
- Dropped any document under 200 characters.
- Decontaminated and deduplicated on the document’s initial 2,000 characters.
- Enforced an aggregate ceiling of 2.2 × 10⁹ characters.
- Routed every 200th document to validation (
val_fraction: 0.005). - Nominal training budget: 800 steps × 128 effective sequences × 2,048 tokens packed = ~210 million processed token positions.
2. Supervised Fine-Tuning Dataset (191,428 Instructions)
Stage 2 converted the domain-adapted base into an instruction-following assistant. We aggregated 215,000 candidate rows across five instruction datasets (shuffled with seed 3407) and filtered down to 191,428 validated pairs:
| Dataset Source | Sample Cap | Instructional Focus |
|---|---|---|
Josephgflowers/Finance-Instruct-500k | 100,000 | Broad financial dialog and instruction following |
sujet-ai/Sujet-Finance-Instruct-177k | 50,000 | Multi-task financial QA and calculations |
gbharti/finance-alpaca | 30,000 | General financial Q&A pairs |
FinGPT/fingpt-sentiment-train | 20,000 | Sentiment classification examples |
FinGPT/fingpt-fiqa_qa | 15,000 | Financial QA from FiQA train splits |
Each row was normalized into ChatML messages. Unparsable rows, duplicates, and contaminated rows were pruned, leaving 1% held out for validation (val_fraction: 0.01).
3. Benchmark Decontamination (13-Gram Shingling)
To prevent test-set leakage, both datasets underwent strict n-gram decontamination against the test splits of the six ChanceFocus/flare-* benchmarks:
- Text is normalized (lowercased, non-alphanumeric characters converted to spaces, whitespace collapsed).
- The benchmark index pools all 13-grams (13 consecutive words) from all test items into a unified global lookup index.
- Any training candidate is purged if its normalized text is an exact match or if 50% or more of its own 13-grams appear in the benchmark index.
- For SFT, non-assistant turns were checked; for CPT, the initial 2,000 characters were verified.
Training Configuration & Hyperparameters
Both stages leveraged parameter-efficient QLoRA on a 4-bit NormalFloat (NF4) quantized backbone with double quantization, executed using TRL’s SFTTrainer. After training, the low-rank adapters were merged back into clean 16-bit BF16 weights.
| Hyperparameter | Stage 1: CPT | Stage 2: SFT |
|---|---|---|
| Starting Checkpoint | Qwen/Qwen3-14B-Base | Merged Finsight CPT weights |
| Tuning Method | QLoRA | QLoRA |
| Backbone Quantization | 4-bit NF4, double quant, BF16 compute | 4-bit NF4, double quant, BF16 compute |
| LoRA Rank (r) | 64 | 32 |
| LoRA Alpha (α) | 128 | 64 |
| LoRA Dropout | 0.05 | 0.0 |
| Target Modules | q, k, v, o, gate, up, down projections | q, k, v, o, gate, up, down projections |
| Packed Sequence Length | 2,048 tokens | 2,048 tokens |
| Per-Device Batch × Grad Acc × GPUs | 4 × 4 × 8 | 8 × 2 × 8 |
| Effective Batch Size | 128 | 128 |
| Training Budget | 800 steps (~210M tokens) | 2 epochs (191,428 examples) |
| Peak Learning Rate | 2e-4 | 2e-4 |
| LR Schedule | Cosine with 2% warmup | Cosine with 3% warmup |
| Optimizer | 8-bit AdamW | 8-bit AdamW |
| Loss Masking | All tokens | All tokens (including prompt turns) |
| Active Duration | ~5 hours | ~2 hours |
Benchmark Evaluation & Results
We evaluated all three checkpoints—Base, Intermediate CPT, and Final CPT-SFT—using lm-evaluation-harness across six FinBen flare-* tasks:
- FinQA & ConvFinQA: Tabular numerical reasoning and conversational financial calculation. Scored with greedy generation up to 256 tokens, credited if the final emitted number matches gold within ±0.001.
- FPB, Headlines & FiQA-SA: 3-class financial news sentiment, gold news classification, and aspect-based sentiment. Scored via option log-likelihood.
- NER: Named Entity Recognition over financial entities, evaluated on set-level F1 score.
Comprehensive Scorecard
| Benchmark | Metric | Qwen3-14B Base | Finsight CPT | Finsight CPT-SFT | Net Δ (Final vs Base) |
|---|---|---|---|---|---|
| FinQA | Accuracy | 0.019 | 0.082 | 0.126 | +0.107 (+563%) |
| ConvFinQA | Accuracy | 0.148 | 0.146 | 0.203 | +0.055 (+37.2%) |
| FPB | Accuracy | 0.812 | 0.790 | 0.794 | −0.018 (−2.2%) |
| Headlines | Accuracy | 0.745 | 0.714 | 0.721 | −0.024 (−3.2%) |
| NER | F1 | 0.238 | 0.150 | 0.180 | −0.058 (−24.4%) |
| FiQA-SA | Accuracy | 0.681 | 0.319 | 0.370 | −0.311 (−45.7%) |
| Six-Task Average | Unweighted mean | 0.441 | 0.367 | 0.399 | −0.042 (−9.5%) |
Note: All tasks were evaluated zero-shot using raw query prompts without applying chat templates, matching standard lm-eval setups.
What the Results Tell Us
- Dramatic Numerical Reasoning Leap: The base model scored a near-zero 1.9% on FinQA. Adding raw financial text (CPT) alone boosted this more than fourfold to 8.2%, and SFT pushed it to 12.6% (a cumulative 6.63× increase). ConvFinQA showed a similar jump to 20.3%. Financial filings contain dense table structures and accounting ratios that the base model rarely prioritized during generic pretraining.
- The SFT Recovery Dynamic: Raw continued pretraining is a blunt instrument. On classification tasks, CPT degraded performance significantly (e.g., FiQA-SA plummeted from 0.681 to 0.319). SFT functioned as a restorative mechanism: it raised scores on every single benchmark compared to the CPT-only model (FiQA-SA recovered by +5.1 pp, NER by +3.0 pp, ConvFinQA by +5.7 pp).
- The Evaluation Format Mismatch: Why did sentiment classification remain below base? During evaluation, tasks were queried with raw prompts and evaluated via choice log-likelihoods without
--apply_chat_template. Because the SFT model was trained with ChatML structure, prompting it with bare text without its learned assistant framing likely penalized its log-likelihood calibration.
Running the Model
The final merged checkpoint is hosted on Hugging Face as standard BF16 Safetensors. It includes Qwen3’s native ChatML template with <think> handling.
Quickstart (Hugging Face Transformers)
Below is an inference script configured to handle stop tokens and thinking blocks cleanly:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "maxint-inc/Finsight-Qwen3-14B-CPT-SFT"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype="auto",
device_map="auto",
)
messages = [
{
"role": "user",
"content": "A company has $120M in current assets, $80M in current liabilities, and $30M in inventory. Calculate and explain its Quick Ratio.",
}
]
# Apply chat template (disable empty thinking blocks to match SFT training layout)
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
# Ensure <|im_end|> is registered as an EOS token to avoid trailing tokens
im_end_id = tokenizer.convert_tokens_to_ids("<|im_end|>")
eos_token_ids = [im_end_id, tokenizer.eos_token_id]
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
eos_token_id=eos_token_ids,
)
new_tokens = output[0, inputs["input_ids"].shape[1] :]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
Colab Smoke Test (Tesla T4 16GB)
We validated the merged checkpoint in a Google Colab notebook on a standard Tesla T4 GPU (14.56 GB usable VRAM):
- Quantization: 4-bit NF4 via
BitsAndBytesConfigwith double quantization. - Footprint: Model loaded at 9.05 GB VRAM, leaving ample headroom for sequence memory (peak allocation ~9.30 GB).
- Sample Test: Given the prompt “A user earns $5,000 monthly and spends $3,200. What is their monthly savings rate? Explain the calculation.”, the model correctly computed the $1,800 monthly savings, derived the 36% savings rate, and produced an accurate step-by-step breakdown.
- Implementation Tip: When running inference, explicitly include
<|im_end|>ineos_token_idas shown in the script above. Because the base model’s defaultgeneration_config.jsononly tracks<|endoftext|>, generation can otherwise spill into trailing tokens before halting.
Limitations & Responsible Use
- Research Artifact: These checkpoints are experimental research artifacts meant for domain-adaptation benchmarking, not production financial advisors.
- No Fiduciary or Safety Alignment: The models have not undergone Reinforcement Learning from Human Feedback (RLHF) or specific financial safety red-teaming. Any calculations or advice must be independently verified.
- Completion vs. Instruction Roles: The CPT checkpoint (
Finsight-Qwen3-14B-CPT) is a base completion model and should not be used as a chat agent without prior instruction tuning.
What’s Next
This experiment gave us concrete data on the behavior of QLoRA-based domain pretraining:
- Chat-Template Re-benchmarking: We plan to re-evaluate the full benchmark suite passing
--apply_chat_templateto isolate format penalty from semantic loss on multiple-choice tasks. - Direct SFT Ablation: Running SFT directly on
Qwen3-14B-Basewithout the CPT stage to measure the exact marginal contribution of raw financial text vs. structured instructions. - Mixture Tuning: Incorporating higher ratios of general reasoning and structured sentiment tasks to prevent classification degradation while preserving tabular calculation gains.
Both checkpoints are now live on Hugging Face. We invite the AI research and financial engineering communities to inspect, test, and build upon them!
- Hugging Face (CPT): maxint-inc/Finsight-Qwen3-14B-CPT
- Hugging Face (CPT → SFT): maxint-inc/Finsight-Qwen3-14B-CPT-SFT
- Google Colab Notebook: Finsight Colab Demo
Stay Informed
Get the latest financial insights, product updates, and wealth-building strategies delivered to your inbox.
No spam, ever. Unsubscribe anytime.