Mistral-7B · Hitchcock Psycho (1960) fine-tune

A hobby fine-tuning experiment. This is a small LoRA adapter for Mistral-7B-Instruct-v0.3, trained with QLoRA on a question-and-answer dataset about Alfred Hitchcock's Psycho (1960). It covers the plot, characters, cast, production and themes.

It was trained and tested on a free Google Colab T4 GPU.

📊 Results viewer Streamlit app — every question, answer and prompt, with a base-vs-fine-tuned comparison
📚 Training data antfr99/hitchcock-psycho-1960-film-dataset (~5.6K Q&A pairs)
🧱 Base model mistralai/Mistral-7B-Instruct-v0.3
🧪 Sister project psycho-mistral-v03-transformed-adapter. It is trained on a deliberately rewritten, fictional version of the same film.

In short: the fine-tuned model gives short, clean answers and gets many well-known facts right. It is also confidently wrong on several plot details. Treat it as a hobby experiment, not a reference source for the film.


What's in this repo

File What it is
adapter_config.json, adapter_model.safetensors The LoRA adapter (~13.6 MB): rank 8, alpha 16, dropout 0.1, applied to q_proj and v_proj.
model-001.safetensors, config.json A 4-bit (bitsandbytes NF4, double-quant) full-model checkpoint, ~4.1 GB.
tokenizer*, special_tokens_map.json, chat_template.jinja Mistral v0.3 tokenizer and chat template.
psycho_keywords.txt Keyword list from an earlier topic-filter experiment.
psycho_inference_OLD_deprecated.py Old inference script. It is kept for reference only; don't use it.

About loading. Because adapter_config.json sits at the top of the repo, recent versions of transformers treat the whole repo as an adapter. Calling AutoModelForCausalLM.from_pretrained("antfr99/mistral-7B-hitchcock-psycho-1960-film") therefore does two things:

  1. It downloads the full Mistral-7B-Instruct-v0.3 base, about 14.5 GB in fp16.
  2. It applies the LoRA on top of that base.

Pass a 4-bit quantization_config so it fits on a T4 (example below).


How to use

# pip install -U transformers accelerate bitsandbytes peft
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig

REPO = "antfr99/mistral-7B-hitchcock-psycho-1960-film"

bnb = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(REPO, quantization_config=bnb, device_map="auto")

# Use the same prompt format as the training data:
prompt = "### Question: Who composed the score for Psycho?\n### Answer:"
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=150, temperature=0.4, do_sample=True,
                     pad_token_id=tok.eos_token_id)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True).strip())

Prompt format. Every training row uses a ### Question: … ### Answer: layout, with no system prompt. That is the format the adapter knows best. The Mistral [INST] … [/INST] chat template also works, and it is what the evaluation below used. But the adapter never saw that format during training.

Colab note. Some Colab images ship an old torchao (0.10). With that version, peft fails while loading the adapter with "Found an incompatible version of torchao", even though torchao isn't used here. Running pip uninstall -y torchao fixes it.

Recommended settings. Temperature 0.35–0.7 and 128–200 new tokens. At 1.5 the answers drift into invention. At 0 with a very small token limit, they just get cut off.


Training

Method QLoRA: 4-bit NF4 base model plus a LoRA adapter (PEFT)
LoRA r=8, lora_alpha=16, lora_dropout=0.1, target modules q_proj, v_proj, no bias
Data hitchcock-psycho-1960-film-dataset: about 5,600 prompt / completion pairs covering plot, characters, cast, production and themes
Hardware Free Google Colab T4 (16 GB)

Example training rows:

### Question: Who directed *Psycho* (1960)?   ### Answer:  Alfred Hitchcock.
### Question: Who wrote the screenplay for *Psycho*?   ### Answer:  Joseph Stefano, based on the 1959 novel by Robert Bloch.

Most completions are one short sentence. The fine-tuned model's short answering style (see below) matches that.


Evaluation: fine-tuned vs. base model, side by side

Setup. 49 questions about the film were each typed once into a Colab/Gradio notebook. The notebook sent each question to two models at the same moment, with the same temperature (0–1.5) and token limit (16–512):

  • Fine-tuned: this adapter on Mistral-7B-Instruct-v0.3, prompted with [INST] … [/INST].
  • Base: Mistral-7B-v0.3, the raw pretrained model with no instruction tuning, prompted with a plain "Question: … / Answer:" text frame.

Both sets of answers are saved to Supabase. You can browse all of them in the Streamlit app, which has a side-by-side viewer and an Explain these results panel.

Answer shape

Fine-tuned Base
Median answer length 17 words 106 words
Answers that ran on into made-up extra Q&A 0 / 49 30 / 49
Answers stuck repeating a line (3+ times) 0 / 49 9 / 49
Answers that used up the token limit 1 / 49* 24 / 49

*This was the one run with a deliberately tiny 16-token limit.

The base model doesn't answer and then stop. It keeps writing the document: after one Q&A pair, the most likely next text is another one. In the first batch of questions the prompts still included list numbers ("Question: 7. …"). The base model picked up that pattern and kept counting, writing questions 8, 9, 10 and so on by itself.

Where the fine-tune gets it right

It gave short, correct answers to many core facts, for example:

  • Anthony Perkins played Norman Bates.
  • Marion steals $40,000.
  • She stops at the Bates Motel, which Norman owns.
  • Lila is Marion's sister.
  • Hitchcock directed the film.
  • Joseph Stefano wrote the screenplay, based on Robert Bloch's 1959 novel.
  • Bernard Herrmann wrote the score.

It also gave sensible short readings of the film's themes: guilt, secrecy, appearance versus reality, and how the past shapes who you are.

Where it's confidently wrong

Question Fine-tuned answer Actually
The woman who steals the money "Caroline (Mary's employer)" Marion Crane
Norman's mother's name "Norman Bates (Voice): Mother." Norma
Hitchcock's rule on late arrivals He let latecomers in (given three times, at different settings) Nobody was admitted after the film started
Who discovers the truth at the end Arbogast Lila (Arbogast is killed before that)
Where Lila finds the mother "Upstairs … behind a false wall" The fruit cellar
Where the opening scene takes place "A remote motel" A hotel room in Phoenix
What the psychiatrist says happened to the mother She "became ill" Norman murdered her

These answers are short and fluent, so they are easier to believe than the base model's long, rambling mistakes. On three of the questions above (the thief's name, the mother's name, and the late-arrival rule), the base model's first line was actually right before it wandered off.

The base model also invents freely:

  • Sam Loomis becomes Marion's "deceased brother".
  • A made-up psychiatrist appears.
  • Ed Gein is described as a woman.
  • Some answers end with leftovers from web pages, such as a Reddit footer, a "Source:" link, or "Back to the question bank".

⚠️ Why this is not a clean "fine-tuned vs. not fine-tuned" test

The two models differ in three ways at once:

  1. Instruction tuning. This adapter sits on the Instruct model. The comparison model is the raw base model. Answering briefly and then stopping mostly comes from instruction tuning, not from the Psycho training.
  2. The prompt. One model got [INST] … [/INST]. The other got a plain text frame.
  3. The adapter itself. This is the only difference that actually comes from the Psycho training.

There are other limits too:

  • Prompt format. The evaluation used the chat template, not the ### Question: format the adapter was trained on. The results may understate what the adapter can do.
  • Grading. Answers are not graded yet. The right/wrong calls above come from a manual read against the film.
  • Sample. Most questions were asked only once, and temperature and token settings changed from question to question.

Next step: compare against Mistral-7B-Instruct-v0.3 with no adapter, using the same prompt. Use the training format, fixed settings and a written grading rubric.


Limitations

  • Not a reliable source on the film. The model gets plot details wrong with full confidence.
  • Knows only one film. It was trained on Psycho (1960) only. Off-topic questions get whatever the base model produces, and nothing filters or refuses them.
  • Small adapter. It is rank 8 and only adapts the attention query and value projections, so it can only shift the base model's behavior so far.
  • English only.

License

MIT for this adapter. The base model, Mistral-7B-Instruct-v0.3, is licensed under Apache-2.0.

Psycho (1960) and its characters belong to their respective rights holders. This is an unofficial fan and educational project.

Downloads last month
94
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for antfr99/mistral-7B-hitchcock-psycho-1960-film

Adapter
(892)
this model

Dataset used to train antfr99/mistral-7B-hitchcock-psycho-1960-film