Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review.
How this is built →
2Distinct papers
7Unique collaborators
2/2Semantic Scholar citation coverage
Publication span: 2026. Corpus fetch span: 2026.
Identity provenance
Provider IDs
- Semantic Scholar:
32733269
ORCID evidence
No valid ORCID is stored.
Observed aliases (1)
- Liudas Panavas (semantic scholar, provider refresh)
Topics and outcomes in this view
Assessment themes
- Human Ai Collab: 2 papers
- Productivity: 2 papers
Claim outcomes
- Organizational Efficiency: 2 papers
- Other: 2 papers
- Error Rate: 2 papers
- Ai Safety And Ethics: 1 paper
- Developer Productivity: 1 paper
- Output Quality: 1 paper
Papers in the Semantic Scholar view
Latest stored Semantic Scholar author observations only. Citation counts below are from the same provider and are not combined with other services.
Scroll the table horizontally to see every column.
| Paper | Author evidence | Date | Provider citations |
|---|---|---|---|
| Post-training on long-horizon office workflows raises a Qwen3.5 agent’s SWE-Bench Pro pass@1 by 5.8 points, with matched-trajectory evidence that the model forms better local goals, builds and preserves task-relevant state, maintains higher-level constraints, and verifies results more often — suggesting long-horizon behavioral training yields domain-general gains.arxiv | Liudas Panavas provider id |
2026-08-03 | 0 |
| A new benchmark finds current language-model agents poorly constrained by long company handbooks: the strongest configuration strictly satisfies all programmatic policy checks in only 36% of tasks, repeatedly ignoring standing rules, corrupting rule details over long horizons, and asserting compliance it did not achieve.arxiv | Liudas Panavas provider id |
2026-07-28 | 0 |
Citation observation summary
Semantic Scholar supplied counts for 2 of 2 papers in this view; 0 are missing. The observed paper counts sum to 0 cumulative citations. This is a coverage summary, not an author score or h-index.