Copied to clipboard!
Module 5 For Educators intermediate 50 min

Evaluation & Feedback with AI

Evidence in. Aligned language out. Judgment, and the signature, stay with you.

For Teachers · Lessons, courses, feedback, family comms

You're viewing the Administrator track. These Teacher exercises are also here. Switch to Teacher in the top bar to focus them.

Draft Student Feedback, Report-Card & IEP Comments

You have a stack of comments to write: report-card narratives, progress reports, IEP or 504 narrative comments, feedback on a common assignment. You want comments that are consistent, specific, and growth-oriented, tied to your rubric, drafted fast and then verified. You do not want generic filler, and you do not want AI inventing evidence about a student you never observed. AI can speed the drafting and the rubric alignment; the evidence, the judgment, and the final language stay yours.

Recommended tool

notebooklm

NotebookLM shines when you upload the rubric and grading criteria as a source and ask it to anchor every comment to a real criterion. It cites the source, which stops AI from inventing evidence about a student. Use Claude as a strong second for longer narrative comments (detailed IEP or conference versions), but only after the rubric and evidence have been mapped.

Bring to the exercise

  • Your rubric or grading criteria (paste in or upload)
  • The redacted scores or observations per student or per category (no student names, no identifying details)
  • A few examples of comments you consider strong (your own, redacted)
  • 1–2 examples of comments you consider weak (if available)
  • Your Educator Context Document (for voice and grade level)
  • Your SIS or gradebook character/word limit for comment boxes (if known)
Step 1

Set the No-Invention Rules Before Pasting Any Student Data

Lock in the rules before AI sees a single score. The comment is about a real student you observed; AI only knows what you tell it.

Confidentiality Guardrail: Read Before Pasting
  • Redact teacher full names. Use a role label like "the teacher."
  • Redact student names and identifying student details.
  • Redact dates of incidents that could re-identify a person.
  • Redact room numbers, period numbers, or schedule details that re-identify the teacher.
  • Redact identifying anecdotes, keep the behavioral evidence, drop the unique detail.
  • Confirm your district has approved the AI tool you are about to paste into.
  • You, not the AI, own the final evaluation. AI cannot invent evidence.

Use the PII Redactor tool to scrub before pasting.

I am a teacher drafting student comments: report-card narratives, progress reports, IEP or 504 narrative comments, and feedback on assignments.

I will provide:
- My rubric or grading criteria
- Redacted scores and observations per student
- Examples of comments I consider strong
- Any SIS or gradebook length limits

Your job is to help me map evidence to criteria and draft comment language.

Hard rules:
- Do not invent evidence. You only know what I give you.
- Do not assume a behavior, score, or growth that is not in what I provide.
- If evidence is missing for a category, call it an "evidence gap" and tell me what is missing. Do not fill it in.
- Use specific, growth-oriented language. Do not exaggerate strengths or weaknesses.
- The teacher (me) makes the final professional judgment and owns every claim.

Confirm you understand these rules before I paste anything.
What good output looks like
  • AI restates the rules in its own words before asking for any data.
  • AI commits to flagging missing evidence as an "evidence gap" rather than inventing it.

Verify before using AI's output

Step 2

Provide the Rubric and Summarize It in Plain Language

Translate your grading criteria into plain language so every comment stays anchored to it.

What to substitute before pasting

  • [PASTE RUBRIC HERE] Your rubric or grading criteria, copy-pasted in full. If it is in a doc, paste the relevant rows.
Here is my rubric or grading criteria:

[PASTE RUBRIC HERE]

Summarize each rubric category in plain language a family could understand.
Then create a short table showing what kind of evidence would support each category. Be specific (e.g. "turned in 8 of 10 assignments," "explained reasoning aloud in discussion," "revised a draft after feedback").
What good output looks like
  • Plain-language summary a parent could read without education jargon.
  • Evidence-type table: specific and observable, not vague.

Verify before using AI's output

Step 3

Paste Redacted Scores and Map Evidence to Categories

Get each student’s evidence sorted into rubric categories, using only what you recorded.

Confidentiality Guardrail: Read Before Pasting
  • Redact teacher full names. Use a role label like "the teacher."
  • Redact student names and identifying student details.
  • Redact dates of incidents that could re-identify a person.
  • Redact room numbers, period numbers, or schedule details that re-identify the teacher.
  • Redact identifying anecdotes, keep the behavioral evidence, drop the unique detail.
  • Confirm your district has approved the AI tool you are about to paste into.
  • You, not the AI, own the final evaluation. AI cannot invent evidence.

Use the PII Redactor tool to scrub before pasting.

What to substitute before pasting

  • [PASTE REDACTED SCORES / OBSERVATIONS HERE] Scores and observations per student or per category: no student names, no identifying details. Run the PII Redactor tool first.
Here are my redacted scores and observations:

[PASTE REDACTED SCORES / OBSERVATIONS HERE]

Sort each piece of evidence into the appropriate rubric category from the previous step.

Use only the evidence I provided. Do not infer beyond what is written.
Where a category has sparse or no evidence, list it as "evidence gap" and note what would have been needed.
Refer to each student as "Student A," "Student B," etc. Never invent a name.
What good output looks like
  • Each rubric category has cited evidence (or is flagged as an evidence gap).
  • No invented evidence: every detail traces back to what you pasted.

Verify before using AI's output

Step 4

Draft Comments in Three Lengths

Get the same evidence shaped for the SIS box, the standard report, and the conference conversation.

Draft comments for each student in three lengths, all anchored to the same evidence:

1. SIS-box version (≤ ~50 words): fits a gradebook or report-card comment box, no markdown
2. Standard version: a short paragraph for a progress report
3. Detailed conference version: a longer narrative for a parent conference or IEP/504 narrative comment

Requirements:
- Use only the evidence mapped in the previous step.
- Be specific. Name the behavior or score, not a generic trait.
- Do not exaggerate. Do not add evidence that is not in the notes.
- Lead with a strength, name one clear growth area, end with a concrete next step.
- Keep the evidence consistent across all three lengths.
What good output looks like
  • Three lengths, same evidence, different depth. The SIS version actually fits in ~50 words.
  • Each comment names a specific behavior or score, not a vague trait like "great student."

Verify before using AI's output

Step 5

Bias and Deficit-Language Check

Catch deficit framing, coded language, and empty praise before a family ever reads it.

Review the draft comments for bias and deficit language. Be skeptical of your own draft.

For each comment, flag and rewrite:
- Deficit framing ("struggles with," "fails to," "lacks") into specific, growth-oriented language tied to a next step
- Coded or character-judging language ("lazy," "disruptive," "unmotivated," "difficult") that describes the student instead of the work
- Vague praise ("great job," "a pleasure to have") that names no evidence
- Any tone difference between comments that the evidence does not justify

Keep every rewrite anchored to the same evidence. Do not soften by inventing positives.
What good output looks like
  • Deficit phrases are rewritten into specific growth language with a next step, not just deleted.
  • AI flags vague praise and replaces it with a named piece of evidence.

Verify before using AI's output

Step 6

Match the Strong Examples, Avoid the Weak Ones

Tune the draft to the comment style you already trust.

What to substitute before pasting

  • [PASTE 2–3 STRONG EXAMPLES] Comments you have written before that you would be proud to send. Redact student names.
  • [PASTE WEAK EXAMPLES] Comments that felt generic or off, so AI knows what to avoid. Redact student names.
Here are comments I consider strong:

[PASTE 2–3 STRONG EXAMPLES]

Here are 1–2 I consider weak:

[PASTE WEAK EXAMPLES]

Tell me what makes the strong ones strong and the weak ones weak (specificity, tone, structure, growth language).
Then revise my drafted comments to match the strong examples and avoid the weak patterns, without changing any of the underlying evidence.
What good output looks like
  • AI names the concrete features that make a strong comment strong (specific evidence, growth framing, family-readable tone).
  • Revised drafts read like your strong examples, with the evidence unchanged.

Verify before using AI's output

Step 7

Final Teacher Sign-Off Checklist

You sign it and it goes home. AI does not.

Final sign-off checklist (you do this manually, not AI):

☐ Is every claim traceable to evidence I actually recorded?
☐ Did I remove every "evidence gap" or fill it with my own real observation?
☐ Is the language specific, not generic ("great student," "needs to try harder")?
☐ Is deficit framing rewritten into growth language with a next step?
☐ Does the SIS-box version fit my gradebook’s length limit?
☐ Would I say this comment to the family’s face?
☐ Is there any student name or identifying detail that must be removed before saving?
☐ Did I copy the final comment into the gradebook/SIS exactly, with no markdown?

If any box is unchecked, do not transfer the comment.
What good output looks like
  • You actually walked through each box before transferring any comment.

Verify before using AI's output

For Admins · PD, evaluation, operations, hard conversations

You're viewing the Teacher track. These Administrator exercises are also here. Switch to Administrator in the top bar to focus them.

Use AI to Support Teacher Evaluation Drafting and Feedback Preparation

You have just completed a classroom observation. You have notes. You have a rubric. You have to translate observation evidence into evaluation language that holds up under review, and you have to be ready for a post-observation conference where the teacher may push back. AI can speed the drafting and rubric alignment, but the judgment, the evidence, and the final language are yours.

Recommended tool

notebooklm

NotebookLM excels when you upload the rubric and observation notes as documents and ask grounded questions. It cites the source, which prevents AI from inventing evidence. Use Claude for the drafting step if you prefer longer-form output, but only after NotebookLM has mapped evidence to rubric categories.

Bring to the exercise

  • Teacher evaluation rubric (paste in or upload)
  • Redacted observation notes (no teacher full name, no student names, no identifying anecdotes)
  • Description of the lesson observed
  • Examples of strong evaluation write-ups (your own, redacted)
  • 1–2 examples of weak evaluation write-ups (if available)
  • [DISTRICT NAME] tone or format expectations
  • Required categories or boxes from the evaluation system
Step 1

Establish the Evaluation Context Before Pasting Anything

Set the rules of engagement before AI sees any evidence.

Confidentiality Guardrail: Read Before Pasting
  • Redact teacher full names. Use a role label like "the teacher."
  • Redact student names and identifying student details.
  • Redact dates of incidents that could re-identify a person.
  • Redact room numbers, period numbers, or schedule details that re-identify the teacher.
  • Redact identifying anecdotes, keep the behavioral evidence, drop the unique detail.
  • Confirm your district has approved the AI tool you are about to paste into.
  • You, not the AI, own the final evaluation. AI cannot invent evidence.

Use the PII Redactor tool to scrub before pasting.

I am a school administrator preparing a teacher evaluation.

I will provide:
- The teacher evaluation rubric
- Redacted observation notes
- Examples of strong evaluation language
- Any required district format expectations

Your job is to help me organize the evidence and draft evaluation language.

Hard rules:
- Do not invent evidence.
- Do not assume anything that is not in the notes.
- If evidence is missing for a rubric category, tell me what is missing. Do not fill it in.
- Use neutral, evidence-based language. Do not exaggerate strengths or weaknesses.
- The administrator (me) makes the final professional judgment.

Confirm you understand these rules before I paste anything.
What good output looks like
  • AI confirms the rules in its own words before asking for evidence.
  • AI commits to flagging missing evidence rather than inventing.

Verify before using AI's output

Step 2

Provide the Rubric

Translate the rubric into plain language and a usable evidence map.

What to substitute before pasting

  • [PASTE RUBRIC HERE] Your district’s rubric, copy-pasted in full. If it’s in a doc, paste the relevant rows.
Here is the teacher evaluation rubric:

[PASTE RUBRIC HERE]

Summarize the rubric categories in plain language.
Then create a table showing what kind of observation evidence would support each category. Be specific (e.g. "student-talk percentage," "wait time after questions," "scaffolding move named in plan").
What good output looks like
  • Plain-language summary that a teacher could understand.
  • Evidence-type table: observable, specific.

Verify before using AI's output

Step 3

Provide Observation Notes

Get evidence sorted into rubric buckets, using only what you saw.

Confidentiality Guardrail: Read Before Pasting
  • Redact teacher full names. Use a role label like "the teacher."
  • Redact student names and identifying student details.
  • Redact dates of incidents that could re-identify a person.
  • Redact room numbers, period numbers, or schedule details that re-identify the teacher.
  • Redact identifying anecdotes, keep the behavioral evidence, drop the unique detail.
  • Confirm your district has approved the AI tool you are about to paste into.
  • You, not the AI, own the final evaluation. AI cannot invent evidence.

Use the PII Redactor tool to scrub before pasting.

What to substitute before pasting

  • [PASTE REDACTED NOTES HERE] Your observation notes: no teacher full name, no student names, no identifying anecdotes. Use the PII Redactor tool first.
Here are my redacted observation notes:

[PASTE REDACTED NOTES HERE]

Sort the evidence into the appropriate rubric categories from the previous step.

Use only the evidence provided. Do not infer beyond what is in the notes.
Where evidence is sparse or missing, list the category as "evidence gap" with a note about what would have been needed.
What good output looks like
  • Each rubric category has cited evidence (or is flagged as a gap).
  • No invented evidence: every quoted detail traces back to your notes.

Verify before using AI's output

Step 4

Identify Evidence Strength and Gaps

Surface where the evaluation will be vulnerable to challenge.

Based on the rubric and observation notes, identify:

- Areas where the evidence is strong (and could survive a teacher pushback or grievance review)
- Areas where the evidence is limited
- Areas where evidence is missing and I need to make a professional judgment
- Questions I should answer myself before drafting the final evaluation
- Anything in the notes that could be interpreted in more than one way
What good output looks like
  • Honest list of weak/strong/missing, not a "looks great!" puff piece.
  • Calls out ambiguity in the notes.

Verify before using AI's output

Step 5

Draft Evaluation Language

Generate the prose, but anchored to evidence only.

Draft teacher evaluation language aligned to the rubric.

Requirements:
- Use only the evidence provided.
- Be specific. Quote or paraphrase the observation notes.
- Do not exaggerate.
- Do not add evidence that is not in the notes.
- Use a professional, neutral tone.
- Make the language easy to transfer into an online evaluation system (concise paragraphs, no markdown).

Tag each sentence with the rubric category it supports.
What good output looks like
  • Every sentence references observable evidence.
  • No hedging adverbs that hide weak evidence ("seems to," "appears to").

Verify before using AI's output

Step 6

Create Version Options

Get the right shape for the evaluation system box and for the conversation.

Create three versions of the evaluation language:

1. Concise version for an online evaluation box (≤ 200 words per category)
2. Detailed version with more explanation (for the formal write-up)
3. Coaching-oriented version that includes specific next steps for teacher growth (used in the post-observation conference)

Keep the evidence consistent across all three versions.
What good output looks like
  • Three versions, same evidence, different lengths and intents.
  • Coaching version names a specific next step, not a vague growth area.

Verify before using AI's output

Step 7

Prepare the Post-Observation Conference

Walk into the conversation with a structure, not a vibe.

Help me prepare for the post-observation conference.

Create:
- Opening statement
- Strengths to name (with the specific evidence)
- Evidence to discuss
- One clear growth area
- Questions to ask the teacher (open-ended, not gotchas)
- Coaching next step
- Closing statement
- Possible teacher pushback and a suggested respectful response (anchor in evidence + rubric)

Keep the tone professional, neutral, and direct.
What good output looks like
  • Open-ended questions outnumber statements.
  • Pushback responses cite evidence, not authority.

Verify before using AI's output

Step 8

Check for Fairness and Accuracy

Stress-test the draft before it leaves your screen.

Review the draft evaluation for fairness and accuracy. Be skeptical of your own draft.

Identify:
- Any claims not supported by evidence
- Any vague language ("appears to," "seems to," "generally")
- Any overly harsh language
- Any overly soft language that hides a problem
- Any missing rubric alignment
- Any places where I need to add administrator judgment that AI cannot supply
What good output looks like
  • AI flags its own weak language.
  • AI names places where you (not it) need to decide.

Verify before using AI's output

Step 9

Final Administrator Review

You sign it. AI does not.

What to substitute before pasting

  • [DISTRICT NAME] Your district name.
Final review checklist (you do this manually, not AI):

☐ Is every claim accurate?
☐ Is every claim evidence-based?
☐ Is every claim aligned to the rubric?
☐ Is the tone appropriate for this teacher?
☐ Is the language professional and neutral?
☐ Does it meet [DISTRICT NAME] format expectations?
☐ Is there any confidential information that needs to be removed before saving?
☐ Are there places where additional administrator judgment is needed?
☐ Could you defend this evaluation in a grievance review?
☐ Have you removed all PII and identifying anecdotes?

If any box is unchecked, do not submit.
What good output looks like
  • You actually walked through each box.

Verify before using AI's output