Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review.
How this is built →
← Authors
Edwin Chen
Provider-ID corpus identity
Auditable corpus entity, not proof of personhood. Provider records and exact ORCID evidence can be incomplete or wrong. Aliases below document source observations; they are not name-based merges. Every count below reflects only papers this corpus has fetched, not this person's complete publication record.
4 Distinct papers
10 Unique collaborators
4/4 Semantic Scholar citation coverage
Publication span: 2026.
Corpus fetch span: 2026.
Identity provenance
Provider IDs
Semantic Scholar : 2444839629
ORCID evidence
No valid ORCID is stored.
Observed aliases (1)
Edwin Chen (semantic scholar, provider refresh)
Topics and outcomes in this view
Assessment themes
Human Ai Collab: 4 papers
Productivity: 4 papers
Papers in the Semantic Scholar view
Latest stored Semantic Scholar author observations only. Citation counts below are from the same provider and are not combined with other services.
Scroll the table horizontally to see every column.
Edwin Chen's distinct papers under the selected provider observation surface.
Paper Author evidence Date Provider citations
Post-training on long-horizon office workflows raises a Qwen3.5 agent’s SWE-Bench Pro pass@1 by 5.8 points, with matched-trajectory evidence that the model forms better local goals, builds and preserves task-relevant state, maintains higher-level constraints, and verifies results more often — suggesting long-horizon behavioral training yields domain-general gains. arxiv
Edwin Chenprovider id
2026-08-03
0
A new benchmark finds current language-model agents poorly constrained by long company handbooks: the strongest configuration strictly satisfies all programmatic policy checks in only 36% of tasks, repeatedly ignoring standing rules, corrupting rule details over long horizons, and asserting compliance it did not achieve. arxiv
Edwin Chenprovider id
2026-07-28
0
Making predictive models first-class tools for LLM agents speeds up work: a pilot system that calls a small pricing model inside an LLM workflow generated priced proposals in under 10 minutes versus multiple hours. The pricing tool—trained on 70 real and human-verified synthetic examples—shows strong in-sample predictive performance, but narrow data and a single pilot limit claims about wider productivity gains. arxiv
Edwin Chenprovider id
2026-02-15
1
LLM-based agents can perform many multi-step workplace tasks but still miss roughly 40% of assignments; failures cluster in tool use, planning and contextual inference, with weaker models failing basic steps and stronger ones stumbling on tasks requiring inference beyond explicit directions. arxiv
Edwin Chenprovider id
2026-01-13
6
Citation observation summary
Semantic Scholar supplied counts for 4 of 4 papers in this view; 0 are missing.
The observed paper counts sum to 7 cumulative citations. This is a coverage summary, not an author score or h-index.