Chat over 1,000 Emails: What Changes When We Add Jev

From individual semantic checks to an application: comparing LLM + full-text search and LLM + Jev on response time, answer completeness and cost.

In our previous article, we tested Jev on operations we usually assign to LLMs: identifying email features and assessing whether a message matches a query. On those tasks, Jev showed a promising combination of speed, cost and accuracy. Classification, in particular, was approximately 4.9 times faster and 8.7 times cheaper in API costs, with slightly better quality.

The next question was how these advantages would translate into a complete application. The system now needs to understand the user’s question, find documents, check conditions, sometimes perform several connected searches, and produce an answer with sources. For this experiment, we chose chat over an email archive and compared it with a conventional approach based on an LLM and full-text search.

Why Email Chat Is a Good Test

An email archive can contain several prices for the same product, a tentative delivery date and its later confirmation. A replacement proposed by the supplier may be rejected by the buyer. The discussion resumes a month later, while the old terms still appear in quoted text.

Three things matter in this kind of chat: responding quickly, considering all relevant documents and keeping costs reasonable. For an internal tool, speed and completeness often matter more than a small difference in query cost: the employee needs an answer so they can get on with their work.

Response time is tied to the volume of analysis. Reading a few well-chosen emails is relatively easy. Checking dozens of similar documents, ruling out refusals and superseded decisions, and reconstructing the connections between them takes more work. The more sources we are willing to check, the better our chance of producing a complete answer, and the more work the system has to do.

We therefore expanded the archive from our first experiment from 100 to 1,000 emails. It contains 77 threads: 990 messages concern the fictional customer Vektor, and another 10 concern Orion. A single product now appears in several projects, similar models are discussed alongside it, and a proposal, its approval and the actual delivery may be recorded in different messages.

Experiment data

All emails were created specifically for this study. The companies, people, products and transactions are fictional; no customer correspondence was used. Below is a message from this synthetic archive.

How Do We Know Which Answer Is Correct?

We first define the transaction histories and facts: products, projects, participants, prices, approvals, deadlines and the sequence of events. We then use these to construct the emails, including initial proposals, clarifications, refusals, old quotations and similar situations. Questions are based on these predefined facts; each has an expected answer and supporting email IDs.

For example, the replacement-bundle question has a known chain: “service case → approved bundle → product → accepted price”. For the question about permission to continue production, we know every email in which permission was actually granted, as well as the documents containing a refusal or an unconfirmed proposal.

The reference answers are stored separately. Indexers and search agents receive the original correspondence and ordinary metadata; they have no access to the reference answers or annotations.

After retrieval, a separate gpt-6-luna call compares the answer with the reference and grades it as complete, partial or incorrect, identifying omissions and errors. For questions asking for a list of emails, we also compare the sets of IDs to find missing or extra matches. Claims are checked against the original messages.

This allows us to assess completeness within the defined scenarios. Automated grading and a small synthetic corpus limit how far we can generalise the result: “12 out of 12” applies to these questions and does not establish universal accuracy across all correspondence. The archive snapshot is fixed as of 18 September 2026.

Indexing: What We Store Before the First Question

Both implementations store the original messages in SQLite. During import, the code parses headers and records ordinary metadata: email and customer IDs, date, thread, reference to the previous message, subject, sender and recipient. Product-code mentions are also extracted locally and stored in emails.metadata.product_ids.

The model then identifies semantic features. Both indexes use the same schema of 15 required boolean fields:

Shared semantic-feature schema: 15 boolean fields
Field in both indexesWhat it checks
has_product_infoNames a product, model, bundle or specific service offered for sale
has_priceStates a price or price range
has_delivery_timeStates a delivery date, period or lead time
has_quantityStates a quantity of goods or services
requests_recipient_actionRequests an action from the recipient in the current message
delivery_is_tentativeThe stated delivery date or lead time is tentative or conditional
has_technical_specsContains a specific technical attribute: power supply, material, pressure, etc.
has_accepted_priceContains a specific price and its current acceptance for a product and order
has_receiving_resultRecords an actual goods-receipt result and quantities
has_product_originStates a product manufacturer or country of origin
has_conformity_certificateContains product-certificate information: ID, validity or status
has_approved_replacement_referenceLinks an approved bundle to a service case, or a change to a project
has_reference_product_mappingExplicitly maps a bundle or change record to a product code
has_actual_commissioning_dateContains the actual date of completed commissioning
has_personnel_tenureNames a person, their role and the period in that role

It is useful to distinguish the presence of information from the values themselves. has_accepted_price=true helps locate an email containing an accepted price. To answer the user, the agent still needs to establish the amount, product, order and applicable pricing terms from the source.

LLM Indexing: A Strict Boolean Schema

The baseline indexer sends gpt-6-luna one email together with descriptions of all 15 features. The response is constrained by JSON Schema and validated with Pydantic. Written out explicitly, the contract looks like this:

from pydantic import BaseModel, ConfigDict, StrictBool

class EmailFeatures(BaseModel):
    model_config = ConfigDict(extra="forbid", strict=True)

    has_product_info: StrictBool
    has_price: StrictBool
    has_delivery_time: StrictBool
    has_quantity: StrictBool
    requests_recipient_action: StrictBool
    delivery_is_tentative: StrictBool
    has_technical_specs: StrictBool
    has_accepted_price: StrictBool
    has_receiving_result: StrictBool
    has_product_origin: StrictBool
    has_conformity_certificate: StrictBool
    has_approved_replacement_reference: StrictBool
    has_reference_product_mapping: StrictBool
    has_actual_commissioning_date: StrictBool
    has_personnel_tenure: StrictBool

In the implementation, the class is created with create_model: each field receives a Field(description=...) from a shared feature dictionary. All fields are required, and extra fields are forbidden. The result is stored both as a complete classification and as individual rows in the flags table, allowing emails to be filtered by feature.

In this implementation, the LLM indexes boolean features. The agent reads prices, manufacturers, countries of origin and relationships between documents from the retrieved originals while preparing its answer. SKUs in ordinary metadata come from local parsing, rather than value extraction by this Pydantic model.

Jev Indexing: Propose a Candidate, Then Check Its Role

With Jev, the same 15 features become Noul questions. The response is a score indicating whether a condition is true; in our experiment, we accept it at a threshold of 0.5.

Values require an additional step. In the interface we use, Jev checks supplied candidates. Code must therefore extract a value from the text first: we wrote regular expressions, then asked Jev whether each extracted string actually played the proposed role in the email.

Email → Regex candidates → Role validation with Jev → Features and values Email Regex candidates Role validation with Jev Features and values

The value index includes 17 fields:

17 value types: regex proposes candidates and Jev checks their roles
GroupValue fields in the Jev index
Products and parties sku, manufacturer, supplier, country_of_origin
Projects and references project, order_id, service_case_id, replacement_bundle_id, change_record_id, quotation_id, certificate_id
People and terms person_name, role, price, quantity, technical_spec, event_date

The classes describing an extracted candidate and the assembled Jev request look like this. payload contains the email and its set of questions, while candidates links each model response to the original value:

from dataclasses import dataclass

@dataclass(frozen=True)
class Candidate:
    field: str       # For example, "country_of_origin"
    value: str       # Normalised value
    raw_value: str   # Exact string from the email
    start: int       # Start of the span in the original
    end: int         # End of the span

@dataclass(frozen=True)
class JevRequest:
    payload: dict
    candidates: dict[str, Candidate]

A single request includes the 15 main checks, additional checks for the presence of a product or a service offered for sale, and questions about every extracted candidate. For each value, we specify its proposed field, context from the email and acceptance criteria. Jev sees the subject and the full message body.

Example: SKUs, a Manufacturer and Italy in Two Roles

Consider email-918. It mentions a standard heat exchanger, another model and different prices in the same message:

For SKUs, we use this pattern:

(?<![\w-])(?P<value>[A-Z]{1,4}-?\d{1,3}[A-Z]{0,2})(?![\w-])

It finds HX-90 and HX-90B. Matching the format still requires a semantic check: a similar string might be a technical designation or a document number. For HX-90, we therefore ask: is this string used as a product code in this email?

A broad supplier pattern also finds Italy before a full stop. The country-of-origin pattern extracts the same string after country of origin is. This produces two candidates with the same value but different proposed fields. Jev checks each in context:

These are actual scores from the indexing log. Italy is retained as the country of origin and rejected as a supplier. Both product codes are confirmed as model mentions: having HX-90B in the index does not mean that its price applies to HX-90.

Storage is divided into three parts:

  • emails.metadata — ordinary metadata and locally extracted product codes;
  • flags — the 15 boolean features, plus their scores for Jev;
  • entities — values accepted by Jev, with their field, score and source span. The complete result, including rejected candidates, is stored in classifications.result.

Here is a shortened representation of the confirmed values from this email:

[
  {
    "field": "sku",
    "value": "HX-90",
    "score": 0.98
  },
  {
    "field": "sku",
    "value": "HX-90B",
    "score": 0.97
  },
  {
    "field": "manufacturer",
    "value": "Aeris",
    "score": 0.98
  },
  {
    "field": "country_of_origin",
    "value": "Italy",
    "score": 0.98
  }
]

This index helps filter documents by specific values and explain where each field came from. However, extraction quality depends on the patterns. If a regular expression misses an unusually formatted SKU or name, Jev never receives that candidate for validation. Extractors need to be extended for new formats and languages. A price mention alone also does not establish which product it applies to: relationships and current terms still need to be checked against the original message.

Time and Cost of Indexing 1,000 Emails

Indexing the full corpus of 1,000 emails
MetricLLMJev + regex
Emails / main decisions 1,000 / 15,000 1,000 / 15,000
Confirmed values Values read when answering 3,753
Successful requests / attempts 1,000 / 1,000 1,000 / 1,050
Sum of API attempt durations 51 min 4 s 26 min 44 s
Wall-clock duration 28.01 s 5 min 35.78 s
Concurrency 128 5
Cost $0.134769 $0.158391 known cost

In this run, the LLM completed indexing faster in wall-clock time. The two tasks have different workloads and concurrency settings: the LLM identifies 15 features, while Jev also validates values. These measurements describe our indexer configurations; comparing model speed under the same workload would require the same execution settings.

Jev confirmed 3,753 values in addition to the 15,000 main boolean decisions. The sum of API attempt durations was approximately 1.91 times lower, but building the index took longer at the chosen concurrency. Indexing cost should be considered separately from the cost of subsequent chat answers.

Four Question Scenarios

We prepared three questions of each type — 12 final questions. They test different ways of working with sources, from retrieving a single fact to checking a new semantic condition.

  1. Find the Source and Answer Immediately
    Example: what power supply, motor rating and wetted casing material were specified for standard P-31 in the Cedar electrical review?
    The answer is in one email: 400 V AC, three phases, 2.2 kW and AISI 316L. The agent needs to find email-152 and distinguish standard P-31 from the P-31H variant. Both agents gave a complete answer.
     
  2. Split the Question into Independent Parts
    Example: how did the accepted P-31 prices change between two orders, who is the manufacturer, and what is the country of origin?
    The two prices and the product’s origin can be retrieved independently and then combined into one answer. The price fell from EUR 1,040 to EUR 980 — a decrease of EUR 60, or approximately 5.77%. The manufacturer is NordWerk, and the country is Finland.
    Both agents answered completely. In this example, the baseline agent took 9.94 s, and the hybrid took 5.33 s.
     
  3. Use a Retrieved Fact to Guide the Next Searc
    Example: which product belongs to the approved replacement bundle for Cedar RMA-Q71, and what price was accepted in SQ-2603?
    First, the service-case decision points to bundle KIT-M9. The bundle register then identifies product SK-8R. Its commercial terms can be found once that relationship is established:
RMA-Q71 → KIT-M9 → SK-8R → EUR 68 RMA-Q71 KIT-M9 SK-8R EUR 68

Both agents reconstructed the chain and gave the price per set, excluding VAT and carriage. Both implementations produced complete answers to all three questions of this type.

4. Check a Condition That Is Not in the Index

The user first asks for an overview of the P-55 discussions. They then follow up: which of these emails record that Vektor actually authorised continued production before the protective interlock had passed its verification?

The emails are already loaded, but this condition was not indexed in advance. We cannot prepare fields for every possible user interest. Usually, candidates are retrieved through full-text search or embeddings, then read and checked. Our baseline uses full-text search; the hybrid checks the new criterion against the original emails with Jev.

The word “approval” alone is insufficient. It appears in permission, refusal, requests for approval and old quotations. The agent needs to establish who gave permission, whether it applies in the current message and whether verification had already been completed.

The first three scenarios describe a sequence of steps; the fourth introduces a new semantic condition. In a real chat, they can be combined.

How the Agents Work

Both agents use gpt-6-luna to plan retrieval and generate an answer with sources. They differ in their prepared indexes and how they check documents.

LLM + FTS searches the 15 features and ordinary metadata. For conditions outside the schema, it uses SQLite FTS5 / BM25 to retrieve text matches, then analyses the original messages with the LLM.

LLM + Jev uses an index of features and confirmed values. Jev identifies the query category, checks candidates and scores their usefulness. When a new condition appears, the LLM formulates it for search_jev, and Jev checks documents within the selected scope.

Search-agent tools
ToolLLM + FTSLLM + Jev
search_by_feature Indexed features and known keys Features and values, followed by Jev validation
search_fts FTS5 / BM25 keyword search —
search_jev — Checks a new criterion against original emails
decompose_query Independent subtasks in parallel Independent subtasks in parallel
refine_search Dependent retrieval using discovered facts Dependent retrieval using discovered facts
read_evidence Reads sources through search and pagination Reads known IDs or a result page
In both agents, decompose_query runs independent subtasks in parallel. refine_search continues retrieval using intermediate facts, such as a bundle code that only became known in the first step. This follow-up retrieval happens inside the agent, without another question to the user. For a new criterion, the hybrid can check the customer’s archive or the document pool saved from the previous turn. The confirmed set of IDs is stored separately from the texts selected for the explanation. The agent can therefore list every match while passing only a small amount of source text to the final LLM.
Tool-Call Diagrams for the Four Scenarios

Simplified routes from the recorded runs for #001, #004 and #007, and the paired check for #010. Green blocks represent tool calls; grey blocks represent internal processing and answer generation. Between calls, the LLM analyses the returned results.

Tool-Call Diagrams for the Four Scenarios

Simplified routes from the recorded runs for #001, #004 and #007, and the paired check for #010. Green blocks represent tool calls; grey blocks represent internal processing and answer generation. Between calls, the LLM analyses the returned results.

001Single source: P-31 specifications

LLM + FTS

LLMPlans retrieval
search_by_featurehas_technical_specs · P-31 · Cedar
LLMAnswers from the retrieved email

LLM + Jev

Jev + indexCategory → candidates → validation → ranking
LLMAnswers from selected sources

004Independent parts: two prices and product origin

LLM + FTS

LLMIdentifies three independent questions
decompose_queryTwo prices and the product passport
search_by_feature × 3Three parallel branches
LLMCombines facts and calculates the change

LLM + Jev

Jev + indexRetrieves initial candidates and pricing facts
LLMRequests missing supporting evidence
read_evidenceEmails containing the prices
search_by_featurehas_product_origin · P-31
LLMCombines results

007Dependent chain: bundle, product, price

LLM + FTS

search_by_featureApproved reference for RMA-Q71
search_by_featureCommercial terms in SQ-2603
search_by_featureMapping from KIT-M9 to the product
LLMConnects documents and answers

LLM + Jev

Jev + indexRetrieves the reference to KIT-M9
search_by_featureProduct mapped to bundle KIT-M9
search_by_featureAccepted price in SQ-2603
read_evidenceReads supporting original messages
LLMProduces an answer with sources

010New condition: permission before interlock verification

LLM + FTS

search_ftsOverview → saved pool
LLMFormulates a full-text query
search_fts × 2Two pages of keyword matches
LLMReads retrieved texts and answers

LLM + Jev

Jev + indexOverview → saved pool
LLMFormulates a new semantic criterion
search_jevsearch_scope=previous_pool · result_mode=list
JevChecks 100 emails and stores 7 matches
LLM + stored IDsExplanation and the complete confirmed list

The Jev agent can receive its initial context automatically, before the first LLM call; this is not an additional function tool. Independent subtasks can use either decompose_query or direct queries with already known keys. A dependent chain can continue through refine_search or another search_by_feature call; example #007 uses direct queries.

Retrieval and Measurement Details

The hybrid processes up to 100 candidates per search. Independent checks are batched into requests of up to 15,000 characters, with up to 20 Jev HTTP requests running concurrently. The model scores source usefulness on a shared Score scale, and the code then sorts the results across batches.

Full source texts totalling up to 3,000 characters are selected for the answer. If a single selected email is longer, it is passed in full. Other sources remain available through read_evidence. Repeated assessments are cached using the criterion, model, history and a hash of the original message.

For the three questions involving a new condition, the initial pool contains exactly 100 emails. The reported completeness applies to the processed pool; candidates beyond this implementation’s limit would require another page or a different traversal method.

Indexing concurrency was 128 for the LLM and 5 for Jev. Ordinary metadata was imported before the timer started. Jev made 1,050 attempts: 50 HTTP 503 responses were retried, and all 1,000 emails were eventually processed. Usage is missing for those 50 attempts, so $0.158391 is the known portion of the cost; the full amount cannot be established from the log.

The full runs of the two agents were executed separately. For question #010, there is also a separate simultaneous run of both agents through the UI. We report those measurements separately. The fourth group’s mean includes the overview and follow-up, excluding the user’s pause; the full series contains 15 turns. Indexing and automated answer grading are excluded from retrieval time and cost.

User-facing time is the elapsed time until the answer is ready. Summed API durations can be higher because independent calls run in parallel. The hybrid’s total token count includes Jev usage; generative LLM tokens are shown in a separate column. Costs are calculated using the rates recorded for the experiment.

Results: Response Time, Completeness and Cost

Below are the results for the 12 scenarios. Each group contains three questions; times are averaged over completed scenarios. For the fourth type, both messages are included: the overview and the follow-up.

Mean response time and quality across three questions per group
ScenarioLLM + FTSLLM + Jev
01Single source 5.30 s2 complete1 partial 3.24 s3 complete
02Independent parts 12.97 s2 complete1 not graded 4.88 s3 complete
03Dependent chain 11.03 s3 complete 10.32 s3 complete
04New condition 20.45 s1 partial2 incorrect 9.34 s3 complete
Complete answer Partial answer Incorrect answer Not graded

The Jev agent received 12 out of 12 complete grades. The baseline had seven complete, two partial and two incorrect answers. No automated grade was recorded for #006, so we do not count that question as an error.

The combined elapsed time across all 15 turns was 149.23 s for LLM + FTS and 83.33 s for LLM + Jev — approximately 1.79 times less for the hybrid. The gain depends on the question: for #009, the hybrid took 17.21 s versus 8.28 s, with both answers graded complete.

How Much Work Remained for the Generative LLM

The hybrid used 155,804 generative LLM tokens, compared with 380,314 for the baseline: a reduction of approximately 59%. However, Jev performs part of the analysis and has its own usage and cost.

Generative LLM tokens and total retrieval cost; totals per group
Scenario LLM tokensbaseline agent LLM tokenshybrid agent USDLLM + FTS USDhybrid
Single source 16,120 13,615 $0.000629 $0.001335
Independent parts 114,753 25,069 $0.004966 $0.002100
Dependent chain 72,569 67,712 $0.002903 $0.003861
New condition 176,872 49,408 $0.008213 $0.016593
All 15 turns 380,314 155,804 $0.016711 $0.023890

The hybrid’s total token count, including Jev, was 545,236. Total retrieval cost in this series was $0.016711 for the baseline and $0.023890 for the hybrid. Broader semantic analysis produced more complete answers and lower combined elapsed time in this experiment, at a higher API cost.

Where Broader Analysis Changes Answer Quality

Return to the question about permission to continue production before interlock verification. In the simultaneous UI run, the baseline reported that there were no matching emails. The Jev agent checked all 100 documents in the selected pool and found seven:

email-134, email-147, email-224, email-287, email-354, email-423, email-490.

The set matched the reference exactly, with no missing or extra IDs.

100emails checked
7matches found
1full email for the explanation

The baseline made three LLM calls. The hybrid made two LLM calls and 35 Jev HTTP requests, checking 100 documents. Across the three new conditions, that amounts to 300 checks of whether an email matches a criterion. Being able to check more possibilities in a reasonable time helps preserve completeness.

The follow-up answer took 7.18 s for the baseline and 6.22 s for the hybrid, at costs of $0.001313 and $0.004689. This is a separate paired measurement and is not included in the full-series table above.

The other two questions involving a new criterion showed a similar pattern: the baseline named two of the seven required emails in #011 and reported no supporting evidence in #012. Jev returned the complete set in both cases. For the simple question #002, the baseline answer was partial because it omitted a carriage condition. Most answers from both agents were complete on independent subtasks and dependent chains.

The Jev agent was also checked against another synthetic archive: seven out of seven questions received complete grades. That check included a new semantic criterion, different products and names, time-period handling, and cases where the documents contained no answer.

Explore One Example of Each Type

Below are recorded answers and their source emails. You can switch scenarios and replay the appearance of each answer using its recorded duration. Questions #001, #004 and #007 use the full runs; #010 uses the separate paired measurement. Viewing the examples makes no new model requests.

What This Means for Email Chat

Our first Jev experiments showed advantages in individual operations. In an application, the result also depends on the index, how candidates are retrieved and the volume of analysis.

In our email chat, Jev made it possible to check more documents against a semantic condition while keeping the generative LLM’s context small. The clearest difference appeared on questions that could not be represented by predefined index fields: checking the full selected pool found supporting evidence that the baseline search missed.

For the practical development of this kind of chat, we would leave value extraction and indexing to the LLM, and use Jev during retrieval. An LLM can extract manufacturers, countries, prices, quantities and document references from free text into a defined structured schema. This reduces dependence on a collection of regular expressions. Values still need to be validated against the contract and stored together with supporting spans.

In the measured baseline indexer, the LLM returned only 15 boolean features. Extending its contract to extract values is the next step in this architecture. The example above illustrates the reason for this division of work: Jev indexing required us to extract candidates separately and check roles such as those proposed for Italy; during retrieval, Jev already receives the original documents and can quickly check the user’s condition.

The combination with embedding-based retrieval is particularly interesting. It can find semantically related documents even when their wording differs from the question. However, semantic similarity does not guarantee the required fact: a request for approval, a refusal, an old proposal and a current authorisation may all appear together. A similar separation between broad candidate retrieval and subsequent assessment is used in the retrieve & rerank approach.

Jev can perform the next step: check retrieved sources against a specific condition, remove unsuitable documents and rank the remaining ones. We can then retrieve a wider candidate pool through embeddings and check more documents before passing supporting evidence to the LLM for its answer. Embeddings broaden coverage, while Jev checks whether sources satisfy the question. This combination could improve completeness within an acceptable response time.

Proposed architecture

Preparation
  1. Emails
  2. LLM: extract values
  3. Metadata + embeddings
Retrieval
  1. Question
  2. Broad candidate pool
  3. Jev: validate and filter
  4. LLM: answer with sources
Indexing and retrieval have different roles: extract values in advance, then check the user’s condition when querying the archive.

We did not measure the combination with embeddings in this experiment; it is a direction for the next implementation. Its completeness also depends on the initial retrieval step: Jev can only check sources that enter the candidate pool.

API costs increased in our series. For an internal chat, this can be a reasonable trade-off if complete answers and shorter waits matter more than a small difference in expenditure. The practical choice depends on the questions in the workflow: which documents need to be considered, how critical an omission would be, and how long the user is willing to wait.

Our conclusion: the LLM prepares structured data and formulates the answer; Jev checks a broad set of sources during retrieval. This division of work lets us use the strengths of both approaches in one email chat.

All 12 Questions and Results

All 12 Questions: Measurements and Full Answers

Tokens in this table are the total API count, including Jev. Indexing and answer grading are excluded. Click a question number to view the full answers. No automated ground-truth grade was recorded for the baseline on #006.

All 12 scenarios; #010–012 include the overview and follow-up
QuestionTime, s
FTS / Jev
Tokens
FTS / Jev
Cost
FTS / Jev
GT
FTS / Jev
#001 4.32 / 4.04 5,232 / 11,934 $0.000195 / $0.000878 Complete/Complete
#002 6.34 / 3.25 5,227 / 6,235 $0.000188 / $0.000221 Partial/Complete
#003 5.23 / 2.43 5,661 / 6,604 $0.000246 / $0.000236 Complete/Complete
#004 9.94 / 5.33 21,557 / 14,516 $0.000891 / $0.000715 Complete/Complete
#005 13.59 / 5.40 37,064 / 11,880 $0.001456 / $0.000577 Complete/Complete
#006 15.39 / 3.90 56,132 / 17,502 $0.002619 / $0.000808 Not graded/Complete
#007 8.07 / 7.20 13,413 / 19,155 $0.000515 / $0.000903 Complete/Complete
#008 16.73 / 6.54 48,584 / 16,453 $0.001869 / $0.000725 Complete/Complete
#009 8.28 / 17.21 10,572 / 52,075 $0.000519 / $0.002234 Complete/Complete
#010 14.63 / 9.01 40,723 / 126,771 $0.002110 / $0.005357 Incorrect/Complete
#011 21.40 / 8.73 34,140 / 126,321 $0.001919 / $0.005371 Partial/Complete
#012 25.31 / 10.28 102,009 / 135,790 $0.004183 / $0.005865 Incorrect/Complete

001What power supply, motor rating and wetted casing material were specified for standard P-31 in the Cedar electrical review?

LLM + FTS

Complete

Standard P-31 was specified for a 400 V AC, three-phase supply, with a 2.2 kW motor and an AISI 316L wetted casing. [email-152]

LLM + Jev

Complete

For the standard P-31 in the Cedar electrical review: 400 V AC, three-phase power supply; 2.2 kW motor; AISI 316L wetted casing. [email-152]

002What was the final accepted net unit price of A-6 under North bay order V-NB-2502?

LLM + FTS

Partial

The final accepted net unit price was EUR 82 per A-6 sensor, excluding VAT. [email-200]

LLM + Jev

Complete

The final accepted price for North bay order V-NB-2502 was EUR 82 net per A-6 sensor, excluding VAT. Carriage was separate. [email-200]

003How many F-12 filters were received in total, quarantined and released for installation under V-EL-2506?

LLM + FTS

Complete

Under V-EL-2506, 30 F-12 filters were received in total: 2 quarantined and 28 released for installation. [email-270]

LLM + Jev

Complete

Under V-EL-2506, 30 F-12 filters were received in total: 2 quarantined and 28 released for installation. [email-270]

004Compare the accepted net P-31 unit prices in Juniper orders V-JU-2501 and V-JU-2601: what changed in EUR and percent, and what manufacturer and country are recorded in its product passport?

LLM + FTS

Complete

P-31’s accepted net unit price decreased from €1,040 in order V-JU-2501 to €980 in V-JU-2601—a reduction of €60 per unit, or about 5.77%. [email-320] [email-400]

The product passport records NordWerk as the manufacturer and Finland as the country of origin. [email-464]

LLM + Jev

Complete

The accepted net P-31 unit price decreased from €1,040 in order V-JU-2501 to €980 in V-JU-2601—a decrease of €60 per unit, or about 5.77%. Both prices are excluding VAT; carriage is separate for V-JU-2601. [email-320] [email-400]

The product passport records NordWerk as the manufacturer and Finland as the country of origin. [email-464]

005What are the accepted quantities and net goods costs for P-24, K-24 and DP-4 under V-WW-2604, their combined goods total, and the DP-4 conformity certificate valid in 2026?

LLM + FTS

Complete

Under order V-WW-2604, the accepted quantities and net goods costs (excluding VAT and carriage) are:

| Product | Accepted quantity | Net goods cost |
|---|---:|---:|
| P-24 | 4 | €2,840 |
| K-24 | 4 | €168 |
| DP-4 | 2 | €760 |
| Combined goods total | | €3,768 |

The DP-4 conformity certificate is CE-DP4-26-017, valid 1 January through 31 December 2026. [email-532] [email-596]

LLM + Jev

Complete

For order V-WW-2604, the accepted quantities and net goods costs are [email-532]:

- P-24 pumps: 4 × EUR 710 = EUR 2,840
- K-24 flange kits: 4 × EUR 42 = EUR 168
- DP-4 drive panels: 2 × EUR 380 = EUR 760

Combined goods total: EUR 3,768, excluding VAT and carriage. The order was accepted; the opening offer was superseded.

The DP-4 conformity certificate is CE-DP4-26-017, valid 1 January through 31 December 2026 for the standard DP-4. [email-596]

006What did the supplier conclude in the Birch valve and filtration reviews about SV-40 and FL-40 for DN 40, 90 °C and 8 bar, with 24 V DC available for the actuator? Compare the stated limits and distinguish the nearby variants.

LLM + FTS

Not graded

The supplier concluded that both standard models meet the stated Birch requirements:

- SV-40: DN 40, rated PN 16 (above the 8 bar operating pressure), maximum liquid temperature 95 °C, and actuator supply 24 V DC—so it meets the 90 °C, 8 bar and electrical requirements. The nearby SV-40H allows 140 °C but requires 230 V AC, unlike the available supply. [email-662]
- FL-40: rated DN 40, PN 10 and 120 °C, so it meets 8 bar and 90 °C. The review also states an AISI 316L body and 100 μm mesh. Nearby FL-40C is limited to 65 °C, so it fails the temperature requirement. [email-682]

LLM + Jev

Complete

The supplier concluded that both standard variants are suitable for the Birch conditions:

- SV-40: DN 40, rated PN 16 and up to 95 °C; its actuator requires 24 V DC. That meets the stated 8 bar, 90 °C and available supply requirements. SV-40H is a different variant: it allows 140 °C but requires 230 V AC, so it does not match the available actuator supply. [email-662]
- FL-40: DN 40, rated PN 10 and up to 120 °C; therefore it meets 8 bar and 90 °C. FL-40C has a 65 °C temperature limit, so it does not meet the 90 °C requirement. The standard FL-40 has an AISI 316L body and 100-micrometre mesh. [email-682]

007Which saleable part belongs to the approved replacement bundle for Cedar service case RMA-Q71, and what accepted net unit price applies in service quotation SQ-2603?

LLM + FTS

Complete

The approved bundle KIT-M9 maps to saleable part SK-8R, the EPDM service seal kit [email-720] [email-784]. In quotation SQ-2603, its accepted net unit price is €68 per kit, excluding VAT and carriage [email-844].

LLM + Jev

Complete

The approved bundle KIT-M9 for Cedar service case RMA-Q71 maps to saleable part SK-8R, the EPDM service seal kit. [email-720] [email-784]

In accepted service quotation SQ-2603, SK-8R is EUR 68 per kit, excluding VAT and carriage. The price is firm for orders placed through September 30, 2026. [email-844]

008Who was Vektor's production superintendent when the Willow recirculation loop was actually commissioned?

LLM + FTS

Complete

Petr Orlov was Vektor’s production superintendent when the Willow recirculation loop was actually commissioned on November 19, 2025. His term ran from September 1 through November 30, 2025. [email-734] [email-900]

LLM + Jev

Complete

Petr Orlov was Vektor’s production superintendent when the Willow recirculation loop was actually commissioned on November 19, 2025. His tenure in the role ran from September 1 through November 30, 2025. [email-734] [email-900]

009For the approved replacement on the Copper cleaning skid, identify the selected product, its manufacturer and country of origin, and its accepted net unit price in the heat-exchanger commercial register.

LLM + FTS

Complete

The approved replacement is HX-90 under change record CHG-K42. Its manufacturer is Aeris, its country of origin is Italy, and its accepted net unit price is EUR 1,530, excluding VAT and carriage. [email-856] [email-918]

LLM + Jev

Complete

The approved replacement for the Copper cleaning skid is HX-90, selected under approved change record CHG-K42; the preliminary HX-90B selection was rejected. [email-802] [email-856]

For HX-90, the manufacturer is Aeris and the country of origin is Italy. Its accepted price in the heat-exchanger commercial register is EUR 1,530 per unit, excluding VAT and carriage, valid through 30 September 2026. This is the register’s accepted product price, not a separately documented case-specific price acceptance. [email-918]

010In those P-55 emails with Vektor, which current messages record that Vektor actually authorised continued production before the protective interlock had passed its verification? Exclude requests, refused proposals, old quotations and cases where verification was already complete.

LLM + FTS

Incorrect

I found no current Vektor message that authorises continued production before the interlock passed verification. The records that might look like authorisations instead say the test had already passed before production was approved [email-143], or explicitly reject continuing before the test and require production to remain stopped [email-350]. Another message says the apparent approval was in a withdrawn old draft and the current instruction was to stop until verification was complete [email-223].

LLM + Jev

Complete

Matching messages (7): [email-134], [email-147], [email-224], [email-287], [email-354], [email-423], [email-490]

Checked condition: The current message records that Vektor actually authorised or approved continued production of P-55 before the protective interlock had passed verification. Exclude messages that merely request authorisation or action, proposals or permissions that Vektor refused, old/obsolete quotations or quoted drafts, and cases where the interlock verification was already completed before the authorisation. Require an affirmative, current Vektor decision to continue production while verification remained incomplete.

011Which of those P-55 messages establish that PotokKomplekt must obtain Vektor's explicit written clearance before substituting a rotor or electrical package, rather than having permission to change it unilaterally?

LLM + FTS

Partial

The clearest evidence is in email-139 (5 April 2024): Vektor says the effective P-55 agreement bars PotokKomplekt from releasing a different rotor or electrical package on its own signature, and requires Vektor Purchasing to send written clearance identifying the changed bill of materials first. It explicitly says silence or packing convenience is not authority to ship the different build.

The same requirement is repeated in email-219 (24 July 2024), which says the restriction is already in force.

There is conflicting earlier wording: emails 138 and 146 say PotokKomplekt may substitute on its own workshop approval and that prior customer consent is not required. Those statements conflict with Vektor’s explicit “effective agreement” clarification in email-139; email-219 later reaffirms that clarification. [email-139] [email-219] [email-138] [email-146]

LLM + Jev

Complete

Matching messages (7): [email-139], [email-219], [email-232], [email-295], [email-359], [email-428], [email-493]

Checked condition: The current email establishes that PotokKomplekt must obtain Vektor's explicit written clearance or approval before substituting a rotor or electrical package, and does not have unilateral permission to make that substitution. Include only messages whose operative content supports this requirement, not merely a proposal, request for approval, or unrelated substitution discussion.

012Which of those P-55 messages commit PotokKomplekt to bearing BOTH removal and reinstallation labour when the defect is caused by its own original assembly, rather than charging Vektor?

LLM + FTS

Incorrect

None of the retrieved P-55 messages establishes that PotokKomplekt bears both removal and reinstallation labour for a defect in its original assembly. Instead, the messages I found say the defect was not in PotokKomplekt’s original assembly, and that Vektor would pay both costs when its installer crossed the wiring. [email-132] [email-290] [email-361] [email-416]

One Oak auxiliary-loop message says PotokKomplekt accepts removal only for a proven P-55 assembly defect; Vektor pays reinstallation. [email-223]

LLM + Jev

Complete

Matching messages (7): [email-142], [email-227], [email-291], [email-363], [email-425], [email-484], [email-495]

Checked condition: The message concerns product P-55 in the Vektor project correspondence and records PotokKomplekt committing that it will bear both removal and reinstallation labor costs if the defect is caused by PotokKomplekt's own original assembly, rather than charging Vektor. Require both removal and reinstallation labor and the stated original-assembly defect condition; exclude requests, proposals without commitment, and messages covering only one labor component.

Insights

LLM Benchmarks September: 225 KI-Models Compared

Gemini pushes into the leading group, Qwen offers a compelling option for local deployments, and DeepSeek shows how much capability is available at a moderate price.

Wissen

TIMETOACT LLM Benchmarks August: 187 AI Models Compared

New models, new rankings: GPT-5.6 Sol, Claude Opus 5, and Grok 4.5 reshuffle the top of the TIMETOACT LLM Benchmarks in August 2026.

CLOUDPILOTS, Google Workspace, G Suite, Google Cloud, GCP, MeisterTask, MindMeister, Freshworks, Freshdesk, Freshsales, Freshservice, Looker, VMware Engine
Produkt

Google Chat

Send direct messages and open group chats with Google Chat. Collaboration is easily designed and efficiently organized.

Blog

Microsoft 365 vs Google Workspace

Google Workspace and Microsoft 365 are two incredibly powerful products. A company needs a central product with which all employees can work together.

Blog 9/20/23

LLM Performance Series: Batching

Beginning with the September Trustbit LLM Benchmarks, we are now giving particular focus to a range of enterprise workloads.

Wissen 6/30/24

LLM-Benchmarks June 2024

This LLM Leaderboard from june 2024 helps to find the best Large Language Model for digital product development.

Wissen 5/30/24

LLM-Benchmarks May 2024

This LLM Leaderboard from may 2024 helps to find the best Large Language Model for digital product development.

Wissen 4/30/24

LLM-Benchmarks April 2024

This LLM Leaderboard from april 2024 helps to find the best Large Language Model for digital product development.

Wissen 7/30/24

LLM-Benchmarks July 2024

This LLM Leaderboard from July 2024 helps to find the best Large Language Model for digital product development.

Insights 3/17/25

LLM Benchmarks: February 2025

Discover the latest insights from our independent LLM benchmarks for February 2025. Find out which large language models performed best.

Wissen 8/30/24

LLM-Benchmarks August 2024

Instead of our general LLM benchmarks, we present the first benchmark of different AI architectures in August.

Wissen 8/30/24

LLM-Benchmarks August 2024

Instead of our general LLM benchmarks, we present the first benchmark of different AI architectures in August.

Insights

LLM Benchmarks March 2025

What's new in the world of LLMs? Find out and read why Google DeepMind managed to surprise us more than once last month.

Wissen

LLM Benchmarks April 2025

Ranking the best performing large language models for digital product development.

Wissen

LLM Benchmarks Summer 2025

Ranking the best performing large language models for digital product development.

Wissen

LLM Benchmarks June 2026

The TIMETOACT LLM Benchmark June 2026 makes it clear: the AI landscape is becoming more competitive and more accessible for businesses.

Wissen

LLM Benchmarks March 2024

The TIMETOACT LLM Benchmarks provide an up-to-date comparison of various Large Language Models to assess their suitability for use in product development.

Wissen

LLM Benchmarks February 2024

The TIMETOACT LLM Benchmarks provide an up-to-date comparison of various Large Language Models to assess their suitability for use in product development.

Blog 9/27/22

Creating solutions and projects in VS code

In this post we are going to create a new Solution containing an F# console project and a test project using the dotnet CLI in Visual Studio Code.

Blog

Answering Business Questions with LLMs

8th place in Enterprise RAG Challenge 2025: Answering Business Questions with LLMs

Bleiben Sie mit dem TIMETOACT GROUP Newsletter auf dem Laufenden!