Launch readiness rarely lives in one place. The state of a launch is spread across status updates, open tickets, test reports, docs, and support notes, and the go/no-go call is a judgment over all of that data. Teams usually settle it in a meeting where everyone reports on their own item.
That meeting looks a lot like a spaceflight launch status check. But Jev can handle that polling.
It is a decision-only model from TypeSafe AI that returns a choice, a score, or a yes/no probability for each question and never writes prose. The declaring stays with a person.
This guide covers what Jev can and can't do in a launch check, and the steps for using it.
Key takeaways
- Jev polls each item and a person decides: It answers closed questions about individual checklist items, and the launch owner declares go or no-go.
- Code handles hard rules and Jev handles judgment: Facts such as CI status or a version bump are read from the source, and Jev reads the text that needs interpreting.
- Uncertain answers go to the item's owner: Probabilities are sorted into a clear go, a clear no-go, and a grey band for a named person, using thresholds set from past launches before the live poll.
- Every result is logged with its evidence: The question, the evidence, the probabilities, and the model version are recorded so the decision can be reviewed.
What Jev can and can't do for a product launch check
A product launch check is a set of closed questions about text, which is the shape Jev handles, using three question types: Choice, Score, and Noul. The limits matter just as much, because they mark where code and people take over.
Here's what Jev can do in a launch check:
- Judge whether evidence shows an item is done: A Noul question returns the probability of yes for a statement such as "the rollback plan has been tested."
- Rate the risk of an open issue: A Score question places it on an ordered scale of 2 to 10 described levels.
- Choose a status: A Choice question picks one option from a list, such as go, hold, or needs review, with a probability for each.
- Check several things per item in one pass: A single request can carry all of an item's questions together.
- Show how sure it is: Each answer carries probabilities, so uncertain items can be routed to the owner.
But there are limitations to it. Here's what Jev can't do:
- Verify that evidence is true: It judges the text it is given, so a note saying "tests passing" is only text to it.
- Guarantee a right answer: It can't return a label outside your list or a malformed answer.
- Pull status from your systems: Evidence has to be gathered and passed in.
- Make the launch decision: It answers questions about individual items and has no view of business stakes or timing.
- Explain an answer: It returns probabilities, not reasons, so a written rationale for an owner needs a general LLM alongside it.
- Read non-text evidence: Screenshots, dashboards, and recordings have to be converted to text first.
- Resist wording written to steer it: Text inside the input can shift its probabilities, which matters here because the people writing status updates have a reason to say "ready."
How to use Jev to check launch readiness
The seven steps below run from writing the questions to recording the decision. The first four set up the poll, and the last three make its results safe to rely on.
Step 1: Turn Each Checklist Item Into a Closed Question
Take every item on the launch checklist and rewrite it as a question with a yes/no answer or a short list of options, with the criteria spelled out. "Is support ready?" is too vague. "Does the evidence show that support has a published macro for each known issue?" is something Jev can answer. Jev reads literally, and specific wording works far better than vague wording, so each question should cover one thing. Write down what "done" means for each item before anything is run, and give every Score level a short written anchor, such as what separates a low-risk open issue from a blocking one, so the scale means the same thing every time.

Step 2: Split the Checks Between Code and Jev
Not every item needs a model. Explicit rules on known fields belong in code, a point a community guide to Jev makes directly. Checks such as CI passing, a version number bumped, or a feature flag existing should be plain code that reads the real system.
Jev is for the checks that need reading text: whether a rollback plan actually describes how to roll back, whether the release notes match what shipped, whether an open bug is a launch blocker. Splitting the work this way also closes the verification gap, because the facts come from the systems and Jev only judges what people wrote.

Step 3: Gather the Evidence as Text for Each Item
For each judgment check, collect the text Jev should read: the ticket summary, the relevant section of the test report, the draft release notes, the support macro, the PR description. Keep it to what the question needs, one item at a time, because extra material adds noise. Anything that isn't text, such as a screenshot or a dashboard, needs a text version first. Keep a link back to the source for every piece of evidence so a reviewer can check the original.
Step 4: Ask Jev the Questions and Pin the Model Version
Send each item's evidence with its questions, several per call where an item has more than one: a Noul question on whether the evidence shows the item is done, a Score for the risk of what remains, and a Choice for go, hold, or needs review. Pin the model version, so a change in results reflects the launch and not a model update.
Step 5: Test the Check on Past Launches and Set Thresholds
Before relying on the poll, run it on a few past launches whose outcomes are known and see whether it would have flagged what went wrong. Set the thresholds and pass/fail lines from those results, and write them down before the live poll so they aren't adjusted after seeing live results. Include status notes written to steer the answer, such as an update that insists an item is "fully ready" without any evidence. The same approach to testing on labeled cases is covered in Jev vs ChatGPT.
Public evidence is still thin. An arXiv audit of the first nine days after Jev's release found its clearest gains in latency and cost, with accuracy gaps on harder tasks, so judge the check by how well it agrees with past outcomes, not by whether its output is well formed.
Step 6: Aggregate With Thresholds and a Grey Band
Each answer is a probability, so the poll needs rules for turning probabilities into outcomes, using the thresholds set from past launches:
- Clear go: The probability is above a high threshold, and the item passes.
- Clear no-go: The probability is below a low threshold, and the item is flagged as a blocker.
- Grey band: Everything in between goes to the item's named owner for a decision.
For yes/no questions, the probability itself is the signal, since Noul answers carry no separate confidence field. Critical items should never pass on a single Jev answer. They need a human sign-off, in line with the common rule that critical gates must pass before a release can be approved, and that sign-off stays even after the check has proven reliable. Automation should widen only for low-risk items.
Step 7: Record Each Result and Have a Person Declare Go or No-Go
Log every result with its evidence, the question asked, the probabilities returned, and the model version, so the decision can be reviewed afterward. Then a named person, the launch owner, reviews the full poll and declares go or no-go. In a launch status check, the poll informs the director's call and doesn't replace it, and the same applies here. Jev answers the item-level questions, and the person who owns the launch decides what the answers add up to. Any override of a no-go gets logged with a reason.

Running the whole poll in one place
The method above stays the same, and the questions, the thresholds, and the final call still belong to the team. What changes is how much of the gathering and running is manual.
Most of the effort in a launch check is not the questions: it is gathering evidence from several tools, running the code checks and the Jev checks together, and keeping the results somewhere the whole launch team can see and correct.
PromptQL covers that part, and in workspaces where Jev has been provisioned it can also run the Jev questions, so a plain-English request can cover the whole poll, as laid out in using Jev without writing code.
- Evidence from several tools in one plan: PromptQL connects to GitHub, Slack, Google Workspace, and other SaaS apps, so the status notes, tickets, and docs a check needs can be pulled together without copying them between tools.
- Code checks and judgment checks together: The model plans, and code runs in a secure sandbox against the real data, so the hard-rule checks read the actual systems and the plan is visible before it runs.
- Access that follows the person: Permissions are enforced at the data layer using the access of the person asking, so people only see the evidence they are allowed to see.
- Criteria corrected in a thread: Owners can be invited into the thread to correct the criteria, then the setup is reused for the next launch or turned into a dashboard.
- A traceable record: Every action is logged and explainable, so the poll can be reviewed after the decision.
Conclusion
A launch is ready when every part of it has been checked by someone accountable for it, and the hard part is collecting those answers consistently from text spread across a team. Splitting the work helps: rules check what can be measured, a model reads what has to be interpreted, and a person owns the final call. Done that way, the go/no-go meeting stops being a round of verbal status updates and becomes a review of recorded answers, with the open items already assigned.
Frequently Asked Questions
Can Jev decide whether a launch should go ahead?
No. Jev answers closed questions about individual items from the text it is given. It has no view of business stakes, timing, or risk appetite, so a named person reviews the full poll and makes the call.
Can Jev verify that tests actually passed?
No. It judges the text it is given, not the system that produced it. Facts such as test results, version numbers, and flag states should be read directly from the source in code, with Jev reserved for items that need interpretation. A well-formed answer also isn't proof of a right one.
What if an owner writes "ready" just to get through?
The check is only as honest as its evidence. Requiring evidence for each item instead of a bare status, testing the poll with status notes written to steer it, and sending critical items to a human sign-off all limit the risk, since text inside the input can shift Jev's probabilities.
Does this work for launches that aren't software releases?
Yes, as long as an item can be phrased as a closed question about text, such as whether a help article covers a known issue or whether an announcement matches the final pricing. Items that depend on non-text evidence need a text version first.