Laya's Prior Art Claim is Absurd

· XTXinverseXTY | John Curcio ·

4 min read Original article ↗

Brief context:

  • Jev came out recently, offering API access to a closed model that answers typed decision questions (developer-specified schemas) in one pass.
  • SalesRLAgent (arXiv, HF repo) is an earlier model by Nandakishor Mukkunnoth that predicts sales-conversion probability from sales conversations.
  • Laya is an open-weights, Jev-compatible alternative, released 3 days after Jev. In its launch post, Mukkunnoth claims SalesRLAgent was prior art for Jev. His claim has since circulated broadly, and many seem confused over precedence.

I dug into SalesRLAgent over the weekend. I found serious errors and nothing uniquely in common with Jev.

Multiple Instances of Data Leakage

  • SalesRLAgent’s train.py includes the eventual conversion outcome in every single observation.1
  • At turn 0, the agent sees an embedding of the entire conversation (including its ending).2
  • conversion_probabilities is initialized with true_probabilities[0], so the ground-truth annotation $q_0$ (see next section) leaks into the state.3

I’ve been guilty of data leakage before and I surely will be again, but it is and forever will be a serious mistake.

Unnecessary Invocation of PPO

The author described it as:

a chess game kinda system for predicting sales conversion probabilities from sales conversations… Then I just trained an RL with PPO, by reducing the dimension using a linear layer and using that to do the final prediction with PPO.

SalesRLAgent seems to be:

  • An MLP over OpenAI text embeddings (plus conversation_metrics, which included the target)
  • Trained on a synthetic dataset of sales conversations (I can’t say which generator revision produced it)
  • RL via PPO?

The model is rewarded based on:

$$ r_t = 1 - |\hat{p}_t - q_t| $$

where $q_t$ is a stored annotation from the synthetic dataset, plus a penalty for being on the wrong side of $0.5$.

If $q_t$ is a latent probability used to generate the synthetic data, then regressing on it is at least a coherent supervised target. From peeking at commit history, that seems not the case here, though I can’t take this as authoritative (generate_dataset.py was deleted and never put back, best I’ve got).

Does the agent learn by intervening in its environment? No, it’s just doing regression: In the training environment, the policy outputs a prediction, which doesn’t affect the progression of the sales conversation.

The author claims that “the guiding brain in my system was always reinforcement learning,” but it’s unclear why PPO is here at all.

Jev

TypeSafe describes Jev as a generally-capable model with a strict, yet generic interface. It guarantees type-safety by restricting the support of the output distribution, and was apparently post-trained with an RL objective that rewards calibration. It may not be calibrated w.r.t. your data but that’s another discussion.

I assume the type-safety guarantee works by constrained decoding, a very old trick4 but admittedly under-exploited. We can’t verify the specific objective or architecture, as it’s not open-source.

Comparison to Jev

From his post, emphasis mine:

They proposed the exact same non-autoregressive decision concept as if it was a brand-new scientific breakthrough.

…

My earlier model used PPO over sequence representations to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0) in vertical sales conversations. Jev generalized parallel sampling using what they called RLCD (Reinforcement Learning for Calibrated Decisions) to output confidence distributions and schema choices horizontally, charging $0.042 per million input tokens with typical response times around 150 ms.

The main point of comparison here appears to be the concept of making decisions based on a non-autoregressive model, using RL. SalesRLAgent’s application of PPO cannot be meaningfully described as RL; I can trivially wrap any scalar regression problem in an RL environment by treating the prediction as an action and the instance loss as a reward.

Point of comparisonSalesRLAgentJev
Support for variable developer schemasNone; SalesRLAgent supported predictions for a single binary outcomeArguably its primary selling point
RL for calibrationReward + penalty not a proper scoring rule, PPO unnecessaryUnknown what RLCD is precisely
Emphasis on salesSales conversations onlyNone
Model architectureMLP over text-embedding-3-largeNot public
Inference costNo hosted API, reliant on OpenAI text embeddings$0.042 / 1M input tokens

Open releases make scrutiny possible, which is one reason they are valuable. But the released SalesRLAgent implementation has disappointing flaws which are immediately obvious upon inspection; I find it hard to believe that this could have inspired Jev.

Its flaws aside, unless the author has nonpublic information about Jev’s training and architecture, they have nothing uniquely in common.