At a glance
| Document type | HR policy template |
|---|---|
| Issued by | Employer |
| Templates included | 3 ready to use versions |
| Download format | Word (.docx) |
| Statutory reference | None cited on this page |
| Last reviewed | 27 August 2026 |
| Maintained by | Engage HR editorial team |
Evaluation form, scorecard and interview notes
Three artefacts from the same process. Confusing them is why panels either capture nothing usable or drown in unread paperwork.
| Interview evaluation form | Role scorecard | Interview notes | |
|---|---|---|---|
| When it is created | After each interview, by that interviewer. | Before any interview, from the job description. | During the interview. |
| What it holds | Ratings against assigned criteria, the evidence, and a recommendation. | The criteria, the scale, and which interviewer assesses what. | What the candidate actually said. |
| Who reads it | The panel at debrief, the hiring manager, HR. | The panel, before they interview. | Mainly the interviewer who wrote them. |
| How long it is kept | For the retention period the organisation has set. | For the life of the opening. | Attached to the evaluation, or discarded on the retention rule. |
| Failure mode | Completed after the debrief, so it records the group view. | Written after the interviews, so it fits the preferred candidate. | Records the interviewer's reactions rather than the candidate's answers. |
How an interview evaluation form is structured
Eight parts. The form should be completable in ten minutes, because a form that takes longer gets completed later, and later means after the debrief.
- Identification. Candidate, requisition, role, band, interview stage, interviewer, date and duration.
- What this interviewer assessed. The criteria assigned to this round, taken from the scorecard rather than chosen by the interviewer.
- The scale. Printed on the form itself, with each point defined. Not a reference to a policy document.
- Criterion blocks. For each criterion: the rating, the evidence observed, and what was not covered.
- Role-specific exercise. Where the round included a task or a case, what was asked and what the candidate produced.
- Candidate questions and disclosures. What the candidate asked, and anything they raised about notice, expectations or availability.
- Recommendation. A single stated position, with the reason in one or two sentences.
- Declarations. Whether the interviewer knows the candidate, and confirmation that the form was completed before the debrief.
Leave out a free text box headed general impression. It is where unstructured judgement goes when the criteria did not capture it, and it is the part of the form that most often records something the organisation would not want to defend.
3 policy templates
Interview Evaluation Form for structured competency evaluation form
The general purpose form for a behavioural round. Note that each criterion asks for evidence and for what was not covered, which is what makes coverage gaps visible at debrief.
[Company Name] INTERVIEW EVALUATION FORM PART A: IDENTIFICATION Candidate: [Candidate Name] Requisition: [Requisition Reference] Role: [Job Title] | Band: [Band or Level] | Location: [Location] Interview stage: [Stage Name] | Format: [In person / Video / Telephone] Interviewer: [Interviewer Name], [Designation] Date: [Date] | Duration: [Duration] PART B: THE SCALE 1 Well below the standard. Could not provide a relevant example, or the example showed the opposite of what the role needs. 2 Below the standard. Examples were relevant but thin, or the candidate described the situation without describing their own contribution. 3 Meets the standard. Clear, relevant examples showing the candidate personally doing what the role requires, at a comparable level of difficulty. 4 Above the standard. As 3, at greater difficulty or scale than the role requires, with reasoning the candidate could explain. 5 Substantially above the standard. Demonstrated at a level that would extend the role as defined. Not assessed. This criterion was not covered in this interview. PART C: CRITERIA ASSESSED IN THIS ROUND Criterion 1: [Criterion Name] What the role needs: [Standard from the scorecard] Rating: [1 / 2 / 3 / 4 / 5 / Not assessed] Evidence observed: [What the candidate described, in enough detail that another reader could form their own view. Record what they did, not how they came across.] Not covered: [Any part of this criterion the interview did not reach] Criterion 2: [Criterion Name] What the role needs: [Standard from the scorecard] Rating: [1 / 2 / 3 / 4 / 5 / Not assessed] Evidence observed: [Description] Not covered: [Description] Criterion 3: [Criterion Name] What the role needs: [Standard from the scorecard] Rating: [1 / 2 / 3 / 4 / 5 / Not assessed] Evidence observed: [Description] Not covered: [Description] [Repeat for each criterion assigned to this round. Do not add criteria that were not assigned.] PART D: CANDIDATE QUESTIONS AND DISCLOSURES What the candidate asked about: [Questions asked] Anything raised about notice period, availability or expectations: [Details, or "nothing raised"] Anything the candidate asked to be followed up: [Details, or "nothing"] PART E: RECOMMENDATION Recommendation: [Advance / Do not advance / Advance with a specific reservation] Reason, in one or two sentences: [Reason, referring to the ratings above] If advancing, what the next round should probe: [Specific area] PART F: DECLARATIONS Do you know this candidate personally or professionally outside this process: [Yes / No] If yes, give details: [Details] I completed this form before the panel debrief: [Yes / No] Signature: ____________________ | Date: [Date]
Interview Evaluation Form for technical or skills round form
Where the round included an exercise. The addition that matters is a record of what was actually asked, so that candidates for the same role are not assessed against different tasks.
[Company Name] INTERVIEW EVALUATION FORM, TECHNICAL ROUND PART A: IDENTIFICATION Candidate: [Candidate Name] Requisition: [Requisition Reference] Role: [Job Title] | Band: [Band or Level] Interviewer: [Interviewer Name], [Designation] Date: [Date] | Duration: [Duration] Format: [Live exercise / Take-home review / Discussion only] PART B: WHAT WAS SET Exercise or problem used: [Exercise Reference from the approved set] Was this the standard exercise for this role: [Yes / No] If no, why a different one was used: [Reason] Time allowed: [Time Allowed] Support or hints given: [What was given, and at what point. Record this, because a candidate who was helped and one who was not have not sat the same assessment.] Environment or tools available: [Description] PART C: ASSESSMENT Scale 1 Well below the standard for [Band or Level] 2 Below the standard 3 Meets the standard for [Band or Level] 4 Above the standard 5 Substantially above the standard Not assessed Criterion 1: Correctness of the solution Rating: [1 / 2 / 3 / 4 / 5 / Not assessed] Evidence: [What worked, what did not, and whether the candidate identified the gaps themselves] Criterion 2: Approach and reasoning Rating: [1 / 2 / 3 / 4 / 5 / Not assessed] Evidence: [How the candidate framed the problem, what they chose to do first, and what tradeoffs they named] Criterion 3: Depth in [Named Skill Area] Rating: [1 / 2 / 3 / 4 / 5 / Not assessed] Evidence: [What they demonstrated and at what level of difficulty] Criterion 4: Response to challenge Rating: [1 / 2 / 3 / 4 / 5 / Not assessed] Evidence: [What happened when a constraint changed or an error was pointed out] Criterion 5: Communication of technical work Rating: [1 / 2 / 3 / 4 / 5 / Not assessed] Evidence: [Whether a non-specialist colleague could have followed the explanation] PART D: WHAT WAS NOT COVERED [Areas of the role's technical requirement this round did not reach, so that a later round can pick them up rather than repeating this ground.] PART E: RECOMMENDATION Recommendation: [Advance / Do not advance / Advance with a specific reservation] Reason: [One or two sentences referring to the ratings above] Band indicated by this round: [Band or Level] If this differs from the band advertised, say why: [Reason] PART F: DECLARATIONS Do you know this candidate personally or professionally outside this process: [Yes / No] If yes, give details: [Details] I completed this form before the panel debrief: [Yes / No] Signature: ____________________ | Date: [Date]
Interview Evaluation Form for panel debrief and decision record
Completed once, after every individual form is in. This is the document that records the decision, and it should show the disagreement rather than smoothing it away.
[Company Name] PANEL DEBRIEF AND HIRING DECISION RECORD PART A: IDENTIFICATION Candidate: [Candidate Name] Requisition: [Requisition Reference] Role: [Job Title] | Band: [Band or Level] | Location: [Location] Hiring manager: [Manager Name] Debrief date: [Date] | Chaired by: [Chair Name] PART B: PANEL AND SUBMISSIONS Interviewer: [Name] | Round: [Stage Name] | Form received on: [Date] | Before debrief: [Yes / No] Interviewer: [Name] | Round: [Stage Name] | Form received on: [Date] | Before debrief: [Yes / No] Interviewer: [Name] | Round: [Stage Name] | Form received on: [Date] | Before debrief: [Yes / No] Any form submitted after the debrief began is recorded here and is given less weight: [Details, or "none"] PART C: RATINGS BY CRITERION Criterion: [Criterion Name] Assessed by: [Interviewer Names] Ratings given: [Ratings] Agreed position: [Agreed Rating] Where the panel differed, what the difference was about: [Description, or "no material difference"] [Repeat for each criterion on the scorecard, including any marked not assessed by every interviewer.] PART D: COVERAGE Criteria not assessed by anyone: [List, or "none"] How this will be resolved: [Additional round / Reference check / Accepted as a known unknown / Not applicable] PART E: RESERVATIONS Reservations raised, and by whom: [Description] What evidence would resolve each: [Description] Whether the panel resolved it or is accepting it: [Description] PART F: DECISION Decision: [Offer / Do not offer / Hold / Additional round] Band for the offer: [Band or Level] Basis for the band: [Reason, referring to the ratings above] Dissent recorded: [Name and position, or "none"] Decision taken by: [Name and Designation] PART G: IF NOT PROCEEDING Reason to be given to the candidate: [Reason, expressed in terms of the role's requirements] Who will communicate it, and by when: [Name] by [Date] Would we consider this candidate for another role: [Yes / No] | If yes, which: [Role] PART H: IF PROCEEDING Offer to be prepared by: [Name] by [Date] Any condition attached to the offer: [Condition, for example completion of background verification] What the first ninety days should focus on, from the reservations above: [Description] Signature of chair: ____________________ | Date: [Date]
What it has to contain
| Element | Why it matters |
|---|---|
| A scale with every point defined in words | Numeric scales without written anchors drift by interviewer, and the drift is invisible. One panellist's four is another's three, and the average of the two records nothing. Definitions printed on the form itself, not referenced from a policy, are what keep the scale stable. |
| An evidence field beside every rating | A rating is a conclusion. The evidence is what allows another reader to reach their own, and what allows the panel to discover at debrief that two people rated the same answer differently for good reasons. A score with no evidence is an impression that has been made harder to question. |
| The criteria assigned to this interviewer, taken from the scorecard | Interviewers who choose their own criteria all choose the ones they find interesting, which produces overlapping coverage of a few areas and none of the rest. Assignment before the interviews is what makes the panel add up to a full assessment. |
| A not assessed option | Without it, interviewers guess at criteria the interview never reached, and the guess is indistinguishable from a judgement. Marking it not assessed is what makes a coverage gap visible in time to do something about it. |
| A single stated recommendation | A form that ends in a paragraph of balanced observation forces the debrief to interpret the interviewer rather than read them. One stated position with a short reason is more useful and more honest, including where the position is to advance with a specific reservation. |
| A declaration of any prior relationship | Interviewers frequently know candidates, and that is not disqualifying. What is damaging is the relationship being unrecorded and surfacing later, at which point every judgement the person made is open to question. |
| Confirmation that the form was completed before the debrief | This single line does more for the integrity of the process than any other part of the form. Scores recorded after a group conversation record the conversation, and the independence the panel structure was meant to provide has already been spent. |
How to write one
- Build the scorecard from the job description. Take the accountabilities and the essential requirements and convert them into four to six criteria the process will assess. Criteria that cannot be traced back to the description are the route by which preference enters as assessment.
- Write the scale before the first interview. Define each point in words, in terms of what a candidate would have to demonstrate. Do this once per role family rather than per opening. An undefined scale is the most common reason panel scores cannot be compared.
- Assign criteria across the rounds. Give each interviewer two or three criteria to assess properly rather than asking everyone to assess everything. Publish the assignment to the panel before the interviews so each person knows what they own and what they can leave alone.
- Brief the panel on the form as well as the role. Ten minutes showing interviewers what the evidence field is for, and what a three looks like against a four, is worth more than any redesign of the form. Most poor evaluation records come from interviewers who were never told what was wanted.
- Collect the forms before the debrief opens. Set the rule that the debrief does not begin until every form is in, and record any form that arrives late. This is the control that keeps the individual assessments independent, which is the entire reason for having a panel.
- Run the debrief criterion by criterion. Work through each criterion, look at where ratings differ, and ask what evidence produced the difference. Panels that open with a general question about impressions converge on the first strong opinion in the room and never recover the detail.
- Record coverage gaps and decide what to do about them. Where no interviewer assessed a criterion, say so in the decision record and choose: another round, a reference check, or an accepted unknown. Gaps that are silently averaged away are the ones that reappear in the first three months.
- Record the dissent. Where a panel member disagreed with the decision, write it down with the reason. It costs nothing, it keeps the record honest, and it is often the most useful thing in the file when the hire is reviewed later.
The scale is the whole design
Most interview scoring problems come from a scale nobody defined. Anchor each point in what a candidate would have to demonstrate, starting from the middle: three means clear, relevant evidence of the candidate personally doing what the role requires. Include a not assessed option, keep to five points, and print the definitions on the form.
Almost every problem with interview scoring traces back to a scale nobody defined.
The default is one to five with no anchors. Interviewers are left to supply their own meaning, and they do. Some read three as adequate for the role. Others read it as a disappointment, since the candidate got through to interview. Some never use one or five at all, compressing everything into the middle. The resulting numbers are then averaged as though they measured the same quantity, and the average is confidently wrong.
The fix costs an hour per role family. Define each point in terms of what a candidate would have to demonstrate. The useful anchor is the middle: three should mean clear, relevant evidence of the candidate personally doing what the role requires, at a comparable level of difficulty. Everything else calibrates from there. Two becomes relevant but thin, or an example where the candidate described a situation without describing their own part in it. Four becomes the same as three at greater difficulty or scale.
Two further design points. Include a not assessed option, because without it interviewers guess at criteria the conversation never reached, and a guess is indistinguishable on the form from a judgement. And keep the scale to five points at most. Longer scales feel more precise and are not, because the extra points have no distinguishable meaning and interviewers use them inconsistently.
Print the definitions on the form. A scale defined in a policy document that interviewers read once during onboarding is a scale that will drift by the third opening.
Evidence is what makes a score reviewable
The evidence box matters more than the rating beside it. A rating is a conclusion and cannot be examined, so two interviewers disagreeing on a number tells you nothing about why. Record what the candidate said and did rather than how they came across. Two or three sentences per criterion is enough.
The most valuable field on the form is not the rating. It is the box beside it.
A rating is a conclusion, and a conclusion on its own cannot be examined. When two interviewers give the same candidate a two and a four on the same criterion, the numbers tell you there is a disagreement and nothing about its nature. With the evidence recorded, the debrief usually discovers something specific. One interviewer asked about a situation the candidate had genuinely handled and the other asked about one they had only observed. Or one accepted a team account where the other pressed for the individual contribution. That is a resolvable difference. Without the evidence it is a difference of opinion between two people, and it resolves according to who is more senior.
What goes in the box matters. The instruction to give interviewers is to record what the candidate said and did, not how they came across. Confident, articulate, seemed nervous and would fit in well are all reactions to a person rather than observations about their work, and a file of them is both useless for deciding and uncomfortable to read back.
There is a further reason to insist on it. Hiring decisions are occasionally questioned, internally or otherwise, and the record is what the organisation has. A file of ratings with no evidence provides nothing to explain why one candidate was preferred. A file of specific observations tied to criteria drawn from the job description explains itself.
The practical concession is length. Two or three sentences per criterion is enough, and asking for more produces forms that get completed the next day, which is to say after the debrief.
Making a panel add up to an assessment
A panel produces overlapping impressions of the same aspect unless someone assigns the coverage. Distribute the four to six criteria across the rounds so each interviewer owns two or three, publish the assignment beforehand, require every form before the debrief opens, and run the debrief criterion by criterion.
Panel interviewing is meant to produce independent assessments of different aspects of a candidate. It usually produces four overlapping impressions of the same aspect, and the reason is that nobody assigned the coverage.
Left to choose, interviewers ask about what interests them. A technical lead asks technical questions, and so does the second technical interviewer. The criteria interviewers find least interesting, which are often the ones about how the person works with others or handles ambiguity, go unasked and are then rated by inference at debrief.
The remedy is assignment. Take the four to six criteria on the scorecard and distribute them across the rounds so each interviewer owns two or three and can go deep on them. Publish the assignment before the interviews. Interviewers find this a relief rather than a constraint, because it tells them what they can safely leave alone.
The second half is sequencing. Every form should be submitted before the debrief opens, and the rule needs to be enforced rather than stated. A panel that discusses first and records afterwards has converted four independent views into one, usually the view of whoever spoke first or ranks highest, and the file will show a unanimity that did not exist.
The debrief itself should run criterion by criterion rather than interviewer by interviewer. Going round the table asking for overall views produces early convergence on a strong opinion. Working through the criteria keeps the discussion on evidence, and it surfaces the coverage gaps, which should be recorded and dealt with rather than averaged away. An unassessed criterion still carries its weight in the average, so the score it produces is an artefact rather than a judgement. It is a question the organisation is deciding to answer after the person starts.
Fairness, records and what the form should not contain
Everything on an evaluation form should trace back to what the role requires. Observations about a candidate's personal circumstances, appearance, accent or family have no place on it, and the general impression box is where they tend to appear. Decide a retention period for unsuccessful candidates' records and apply it consistently.
An evaluation form is a record about an identified individual, usually one who does not work for the organisation and may never do so. That has consequences for what belongs in it and how long it is kept.
On content, the discipline is that everything on the form should trace back to what the role requires. Criteria drawn from the job description, evidence about what the candidate has done, and a recommendation grounded in the ratings. Observations about a candidate's personal circumstances, their appearance, their accent or their family situation have no place on the form. They are not relevant to whether the person can do the job, and they are difficult to explain if the file is ever read by anyone outside the panel.
The general impression box is where these tend to appear, which is the main argument for removing it. When interviewers consistently want to write something that the criteria do not accommodate, the right response is to examine whether a criterion is missing rather than to provide an unstructured space for whatever they were thinking.
On retention, an organisation collecting evaluation records should know how long it keeps them, where they are held, who can read them, and what happens to the records of candidates who were not hired. Many organisations have never decided any of this, and unsuccessful candidates' files accumulate indefinitely in whatever system the recruitment team was using at the time. Data protection obligations in India are developing, and the specific requirements applicable to a given organisation should be confirmed against the current position rather than assumed. What is clear regardless is that a stated retention period, applied consistently, is a better place to be than an unbounded archive nobody has looked at.
A related point on transparency. Where a candidate asks why they were not selected, a process built on criteria and evidence can give a real answer expressed in terms of the role's requirements. That is better for the candidate, better for the organisation's reputation with people it may want to hire later, and only possible if the form was completed properly in the first place.
Common mistakes
| Mistake | Why it causes trouble | What to do instead |
|---|---|---|
| An unlabelled numeric scale | Interviewers are asked to rate from one to five with no definitions. Some treat three as adequate, others as disappointing. The scores are averaged as though they measured the same thing, and the average is meaningless without showing it. | Write a definition for each point in terms of what the candidate would have to demonstrate, and print the definitions on the form. |
| Ratings without evidence | The form records a two on communication and nothing else. At debrief the panel cannot tell whether the candidate was unclear, was quiet, or simply disagreed with the interviewer, and the number cannot be challenged because there is nothing behind it. | Make the evidence field mandatory and brief interviewers to record what the candidate said and did rather than how they came across. |
| Forms completed after the debrief | The panel discusses the candidate, forms a collective view, and then everyone writes it down. The four independent assessments the process was designed to produce have become one, and the file looks like a consensus that never existed. | Require submission before the debrief opens, record any late form, and give it less weight in the decision. |
| Everyone assessing everything | Four interviewers each ask about the same two areas because those are the ones they find interesting. Half the criteria are covered four times and half are not covered at all, but the total volume of paperwork disguises the gap. | Assign two or three criteria to each round from the scorecard, and publish the assignment to the panel in advance. |
| A general impression box | The field collects everything the criteria did not, which is precisely the unstructured judgement the form exists to discipline. It is also the part of the record most likely to contain something the organisation could not defend if the file were ever read by someone else. | Delete it. Where interviewers are consistently writing something important in it, that is a signal a criterion is missing from the scorecard. |
| Criteria invented at interview stage | An interviewer decides the role needs a quality the description never mentioned, rates against it, and the panel treats it as a legitimate finding. Whatever the intention, this is how a preference becomes an assessment. | Restrict the form to the criteria on the scorecard. Where an interviewer believes a criterion is missing, it should be added to the scorecard for all candidates rather than applied to one. |
Frequently asked questions
What should an interview evaluation form include?
Identification of the candidate, role and interviewer, and the criteria assigned to that round from the scorecard. A scale with each point defined in words, a rating and an evidence field for each criterion, and a not assessed option. Then a single recommendation with a short reason, and declarations covering any prior relationship and confirming the form was completed before the debrief.
What rating scale works best for interview evaluation?
Three to five points, each defined in words. Anchor the middle point as clear evidence of the candidate personally doing what the role requires at a comparable level of difficulty, and calibrate the rest from there. Longer scales feel more precise but the extra points have no distinguishable meaning, so interviewers use them inconsistently.
Should interviewers complete the form before or after the panel discussion?
Before, and the rule needs enforcing. Scores recorded after a group conversation record the conversation rather than the interview, which removes the independence the panel structure existed to provide. Record any form that arrives late and give it less weight.
How many criteria should an interview evaluation form cover?
Four to six across the whole process, with each interviewer assessing two or three properly. Asking everyone to assess everything produces repeated coverage of the areas interviewers find interesting and no coverage of the rest.
Should the form include a general impression section?
Better not. It collects exactly the unstructured judgement the form exists to discipline, and it is the part of the record most likely to contain something the organisation could not defend. Where interviewers consistently want to write something there, that is a signal that a criterion is missing from the scorecard.
How long should interview evaluation records be kept?
For a stated period the organisation has decided on, applied consistently to successful and unsuccessful candidates alike. Data protection requirements in India are developing and the position applicable to a particular organisation should be confirmed against the current law, but an unbounded archive of candidate assessments is a poor place to be regardless.
What is the difference between an evaluation form and a scorecard?
The scorecard is written before any interview, from the job description, and defines the criteria, the scale and who assesses what. The evaluation form is completed after each interview by that interviewer and records ratings, evidence and a recommendation against the criteria the scorecard assigned them.
How should a panel resolve disagreement between interviewers?
By working through the criteria one at a time and asking what evidence produced the difference, rather than by asking each person for an overall view. Differences usually turn out to be about what was asked or what was accepted as an example, which is resolvable. Where the disagreement stands, record the dissent in the decision.
Interview records in Engage
Engage holds the scorecard against the requisition, so each interviewer sees only the criteria assigned to their round and the debrief cannot open until every form is in. Ratings, evidence and the decision record stay attached to the candidate for the retention period the organisation has set, and drop out of the system when it expires rather than accumulating in whoever's drive the process ran through.
Book a demo