Step 1: Gather the actual work, not opinions about it
Before you name a single competency, collect artifacts of the job being done well. For a customer support role, pull twenty resolved tickets from your two strongest agents — including the messy ones with an angry opener or a policy gray area. For inside sales, pull call recordings, follow-up emails, and CRM notes from reps who consistently convert. You are building a corpus of what good actually looks like, in the words and formats the job really uses.
Resist the urge to start from the job description. Job ads describe an idealized person; artifacts describe real performance. When a support lead says the role needs 'strong communication,' the tickets will show you what that means in practice: acknowledging the customer before proposing a fix, quoting the exact policy line, closing with a clear next step. The artifacts turn vague traits into observable material you can design against.
Interview the people closest to the work last, not first. Ask a top performer to walk you through one hard ticket or one stalled deal: what did they notice, what did they decide, what would a weaker colleague have done instead? That contrast — strong versus average on the same input — is the raw material for the next step.
- Ten to twenty real work products per role: tickets, call recordings, emails, proposals, QA reviews.
- At least a few difficult cases — escalations, exceptions, objections — not just routine wins.
- One walkthrough interview per top performer, focused on decisions rather than duties.
- Notes on what average performers do differently on the same kind of input.
Step 2: Extract 4–6 competencies that differentiate performance
Read through your corpus and mark every moment where a strong performer did something an average one would not. Cluster those moments into themes. You are not listing everything the job involves — you are isolating the four to six competencies that separate a great hire from an acceptable one. Everything a new hire learns in the first two weeks, like product facts or tool navigation, gets cut.
For a support role, the differentiators usually land near judgment within policy, de-escalation in writing, diagnostic questioning, and follow-through. For inside sales, they cluster around discovery questioning, objection handling, pipeline discipline, and written follow-up. Your list will differ, and it should — the point of working from artifacts is that the blueprint reflects your customers, your policies, and your motion.
If your list runs past six, you are probably mixing differentiators with baseline requirements. Ask of each candidate competency: did the artifacts show strong and average performers doing this differently? If everyone does it about the same way, it does not belong in the blueprint, however important it sounds — it belongs in onboarding, or in a simple screening requirement.
Step 3: Define observable behaviors for each competency
A competency name is not assessable; a behavior is. For each of your four to six competencies, write two to four statements describing what an evaluator could literally see or read in a candidate's response. The test for each statement: could two evaluators look at the same work product and agree whether the behavior happened? 'Shows empathy' fails that test. 'Acknowledges the customer's frustration before proposing a fix' passes it.
Pull the language directly from your artifacts. If your best support agents consistently restate the customer's problem in one sentence before answering, that exact move becomes a behavior. If your best inside sales reps always confirm budget and timeline before proposing a demo, write that down verbatim. Borrowed behavior lists from generic frameworks will not survive contact with your rubric later — behaviors grounded in your own corpus will.
- Judgment within policy: applies the refund rule correctly, and names a fair exception when the rule produces an absurd outcome.
- De-escalation in writing: acknowledges the frustration, avoids defensive language, and commits to a specific next step with a time.
- Discovery questioning: asks an open question about the prospect's current process before pitching anything.
- Follow-through: summarizes agreed actions in writing and flags what needs a handoff.
Step 4: Map each competency to a task type
Now decide how each competency will be observed. The rule is fidelity: choose the task format closest to how the behavior shows up on the job. Written de-escalation is observed in a ticket simulation, not a personality question. Live objection handling is observed in an audio or video response to a recorded prospect objection, not a multiple-choice quiz about sales methodology.
One task can cover more than one competency, and usually should. A single support ticket with an angry customer and a policy edge case can surface judgment, tone, and written clarity at once. Aim for a sequence of two to four tasks that together touch every competency at least once, with your most decision-critical competencies observed in more than one task.
Write the mapping down as a table in the blueprint: competency, behaviors, task type, and which task in the sequence covers it. This table is the contract for everything downstream — when you write the actual tasks, and later the rubric, every design choice traces back to a row here, and gaps or duplicated coverage become visible at a glance.
- Judgment within policy → ticket simulation with a gray-area refund request.
- De-escalation in writing → written reply to a frustrated customer message.
- Discovery and objection handling → audio response to a recorded prospect scenario.
- Pipeline discipline → short case study prioritizing a mock CRM queue.
- Written follow-up → file upload of a post-call summary email.
Step 5: Weight what matters before you see any candidates
Not all competencies deserve equal weight, and deciding weights after you have seen responses invites motivated reasoning. Set them now, while the blueprint is still abstract. Ask the hiring manager: if a candidate were strong on everything except one competency, which gap would you refuse to hire around? That competency gets the heaviest weight. Which gap could coaching close in a quarter? That one weighs less.
For a support team drowning in escalations, judgment within policy might carry twice the weight of troubleshooting speed. For an inside sales team with strong leads but weak conversion, discovery questioning might dominate. Record the weights and one sentence of rationale for each — the rationale is what lets you defend the blueprint when a stakeholder later asks why their favorite trait counts for less.
Step 6: Version the blueprint and reuse it across reqs
Treat the finished blueprint like a spec, not a document that quietly drifts. Give it a version number, note the date and the artifacts it was derived from, and freeze it. Every assessment, rubric, and decision for this role now references blueprint v1 — which means candidates assessed in March and candidates assessed in September were measured against the same standard, and you can prove it.
When the role genuinely changes — a new product line shifts what support judgment means, or the sales motion moves upmarket — revise deliberately. Gather fresh artifacts, rerun the extraction, and publish v2 with a short changelog. Reqs in flight finish on the version they started on. This is what makes a blueprint reusable: not that it never changes, but that every change is explicit, dated, and explainable.
The payoff compounds. The second req for the role skips straight to task selection. Adjacent roles — support to success, SDR to AE — start from a copy of the nearest blueprint rather than a blank page. And when anyone asks why a candidate was assessed the way they were, the answer is a document, not a memory.
Key takeaways
- Start from artifacts of real work — tickets, calls, deliverables — never from the job ad.
- Keep only the 4–6 competencies where strong and average performers visibly differ.
- Write behaviors two evaluators could independently confirm in a work product, using language from your own corpus.
- Map every competency to the highest-fidelity task type and set weights before seeing any candidate.
- Version the blueprint, freeze it per req, and revise with a changelog when the role genuinely changes.